用关键词检索提升濒危语言翻译质量
Transcending Language Boundaries: Harnessing LLMs for Low-Resource Language Translation
- 聚焦关键术语,从已有数据中检索对应翻译例句
- 在切罗基语、藏语和满语上显著提升翻译准确率
- 适合关注小语种保护与AI公平性的研究者
大型语言模型在诸多任务中表现卓越,但在低资源语言翻译,尤其是向这些语言的翻译方面仍缺乏充分探索。这一差距阻碍了少数族群的文化传承与发展。为此,本文提出一种基于检索的新方法,通过聚焦关键词并从现有数据中检索对应翻译示例,来提升低资源语言的翻译质量。我们在英语转切罗基语、藏语和满语三类低资源语言上进行了实验。与GPT-4o和LLaMA 3.1 405B的零样本表现相比,该方法在词级准确率和整体语义理解上均有明显提升,证明了其在有效利用现有资源方面的潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable success across a wide range of tasks and domains. However, their performance in low-resource language translation, particularly when translating into these languages, remains underexplored. This gap poses significant challenges, as linguistic barriers hinder the cultural preservation and development of minority communities. To address this issue, this paper introduces a novel retrieval-based method that enhances translation quality for low-resource languages by focusing on key terms, which involves translating keywords and retrieving corresponding examples from existing data. To evaluate the effectiveness of this method, we conducted experiments translating from English into three low-resource languages: Cherokee, a critically endangered indigenous language of North America; Tibetan, a historically and culturally significant language in Asia; and Manchu, a language with few remaining speakers. Our comparison with the zero-shot performance of GPT-4o and LLaMA 3.1 405B, highlights the significant challenges these models face when translating into low-resource languages. In contrast, our retrieval-based method shows promise in improving both word-level accuracy and overall semantic understanding by leveraging existing resources more effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。