为印度东北部的科克博罗克语打造高质机器翻译系统
Towards High-Quality Machine Translation for Kokborok: A Low-Resource Tibeto-Burman Language of Northeast India

- 基于多源语料微调NLLB模型,构建低资源语言翻译系统
- 在测试集上达到17.30和38.56的BLEU分数,显著提升
- 首次为科克博罗克语引入专用语言标记,适合濒危语言研究者
我们提出KokborokMT,一个针对科克博罗克语(ISO 639-3)的高质量神经机器翻译系统。该语言是印度特里普拉邦的主要语言,约有150万使用者,但长期缺乏自然语言处理资源,此前的翻译系统仅基于小规模圣经语料训练,BLEU得分低于7。本研究将NLLB-200-distilled-600M模型在包含36,052句对的多源平行语料上进行微调:包括9,284条来自SMOL数据集的专业翻译句、1,769条来自WMT共享任务的圣经领域句子,以及24,999条通过Gemini Flash从Tatoeba英文句子生成的合成回译句。我们在NLLB框架中引入了科克博罗克语的新语言标记。最佳系统在保留测试集上分别取得17.30和38.56的BLEU分数,远超先前结果。三人人工评估显示平均适切性为3.74/5,流畅性为3.70/5,评估者间一致性较高。
原文摘要 · Abstract (English)
We present KokborokMT, a high-quality neural machine translation (NMT) system for Kokborok (ISO 639-3), a Tibeto-Burman language spoken primarily in Tripura, India with approximately 1.5 million speakers. Despite its status as an official language of Tripura, Kokborok has remained severely under-resourced in the NLP community, with prior machine translation attempts limited to systems trained on small Bible-derived corpora achieving BLEU scores below 7. We fine-tune the NLLB-200-distilled-600M model on a multi-source parallel corpus comprising 36,052 sentence pairs: 9,284 professionally translated sentences from the SMOL dataset, 1,769 Bible-domain sentences from WMT shared task data, and 24,999 synthetic back-translated pairs generated via Gemini Flash from Tatoeba English source sentences. We introduce as a new language token for Kokborok in the NLLB framework. Our best system achieves BLEU scores of 17.30 and 38.56 on held-out test sets, representing substantial improvements over prior published results. Human evaluation by three annotators yields mean adequacy of 3.74/5 and fluency of 3.70/5, with substantial agreement between trained evaluators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。