用大模型和检索增强提升客家话翻译质量,尤其擅长文化专有词。
Enhancing Low-Resource Minority Language Translation with LLMs and Retrieval-Augmented Generation for Cultural Nuances
- 结合检索与大模型生成,动态补充文化术语
- 最佳模型在客家话翻译上达31% BLEU,显著优于纯词典方法
- 适合低资源语言研究者,强调社区协作与文化保真
本研究探讨将大语言模型(LLMs)与检索增强生成(RAG)结合以应对低资源语言翻译挑战。在客家话翻译任务中,仅使用词典的模型BLEU仅为12%,而采用Gemini 2.0的RAG模型达到31%。表现最佳的Model 4融合检索与高级语言建模,在专业及文化特有词汇的覆盖度与语法连贯性上均有提升。两阶段方法(Model 3)先用词典生成,再由Gemini 2.0修正,取得26% BLEU,体现迭代优化价值。静态词典难以处理上下文敏感内容,暴露出依赖预定义资源的局限性。研究强调需构建精选资源、融入领域知识,并与本地社区开展伦理合作,为提升翻译准确率与流畅性、支持文化传承提供可行框架。
原文摘要 · Abstract (English)
This study investigates the challenges of translating low-resource languages by integrating Large Language Models (LLMs) with Retrieval-Augmented Generation (RAG). Various model configurations were tested on Hakka translations, with BLEU scores ranging from 12% (dictionary-only) to 31% (RAG with Gemini 2.0). The best-performing model (Model 4) combined retrieval and advanced language modeling, improving lexical coverage, particularly for specialized or culturally nuanced terms, and enhancing grammatical coherence. A two-stage method (Model 3) using dictionary outputs refined by Gemini 2.0 achieved a BLEU score of 26%, highlighting iterative correction's value and the challenges of domain-specific expressions. Static dictionary-based approaches struggled with context-sensitive content, demonstrating the limitations of relying solely on predefined resources. These results emphasize the need for curated resources, domain knowledge, and ethical collaboration with local communities, offering a framework that improves translation accuracy and fluency while supporting cultural preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。