arXiv:2412.05184cs.CLcs.AI2024-12被引 5

用检索增强与高效微调,提升濒危语言克丘亚语翻译质量

QueEn: A Large Language Model for Quechua-English Translation

  • 结合检索增强生成与低秩适配微调,利用外部语言资源
  • 克丘亚语-英语翻译BLEU达17.6,远超基线模型的1.5
  • 适合关注濒危语言保护与低资源翻译的研究者

近期研究表明,大语言模型在自然语言处理中表现强大,推动了计算语言学诸多进展。然而,由于训练数据有限且难以理解文化细节,这些模型在低资源语言上面临挑战。本文提出QueEn,一种面向克丘亚语-英语翻译的新方法,融合检索增强生成(RAG)与参数高效微调技术。该方法通过RAG利用外部语言资源,并采用低秩适配(LoRA)实现高效模型调整。实验结果表明,所提方法显著优于基线模型,克丘亚语-英语翻译的BLEU分数达到17.6,而标准GPT模型仅为1.5。RAG与微调的结合使系统在应对低资源语言翻译挑战的同时保持计算效率。本工作为通过先进语言技术保护濒危语言提供了重要贡献。

原文摘要 · Abstract (English)

Recent studies show that large language models (LLMs) are powerful tools for working with natural language, bringing advances in many areas of computational linguistics. However, these models face challenges when applied to low-resource languages due to limited training data and difficulty in understanding cultural nuances. In this paper, we propose QueEn, a novel approach for Quechua-English translation that combines Retrieval-Augmented Generation (RAG) with parameter-efficient fine-tuning techniques. Our method leverages external linguistic resources through RAG and uses Low-Rank Adaptation (LoRA) for efficient model adaptation. Experimental results show that our approach substantially exceeds baseline models, with a BLEU score of 17.6 compared to 1.5 for standard GPT models. The integration of RAG with fine-tuning allows our system to address the challenges of low-resource language translation while maintaining computational efficiency. This work contributes to the broader goal of preserving endangered languages through advanced language technologies.

机器翻译低资源语言克丘亚语RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。