为濒危的丁古尔语-英语翻译构建首个低资源模型,性能达基线水平。
Neural Machine Translation for Low-Resource Tangkhul--English
- 用ByT5-large微调3.8万句平行语料,实现端到端翻译
- 在3856句测试集上达到BLEU 39.97,超越多数低资源模型
- 针对拉丁字母变音符号难题提出可复用的处理方案
我们针对丁古尔语-英语(nmf-en)这一严重低资源语言对开展机器翻译研究。丁古尔语是印度曼尼普尔邦主要使用的藏缅语系语言,此前几乎无自然语言处理基础。本文提出两个系统:(1) 基于ByT5-large,在38,336句平行语料上微调的主系统;(2) 基于mT5-small的对比系统,同样在相同语料上训练。主系统在3,856句的保留测试集上取得BLEU 39.97、chrF++ 58.07、BERTScore F1 0.8104、COMET (wmt22-comet-da) 0.7302的成绩。我们还讨论了丁古尔语拉丁字母变音符号带来的书写挑战,以及训练语料存在的领域偏差(涵盖圣经文本、故事与对话数据),并指出未来可通过数据多样化和领域适配进一步提升性能。
原文摘要 · Abstract (English)
We present a study on low-resource machine translation for the Tangkhul-English (nmf-en) language pair. Tangkhul is a severely under-resourced Tibeto-Burman language spoken primarily in Manipur, India, with virtually no prior natural language processing infrastructure. We describe two systems: (1) a primary system based on ByT5-large fine-tuned on 38,336 Tangkhul-English parallel sentence pairs, and (2) a contrastive system based on mT5-small fine-tuned on the same corpus. Our primary ByT5-large system achieves a corpus BLEU score of 39.97, chrF++ of 58.07, BERTScore F1 of 0.8104, and COMET (wmt22-comet-da) of 0.7302 on a held-out test set of 3,856 sentences. We further discuss the orthographic challenges specific to Tangkhul's Latin-script diacritics, the domain bias of our training corpus (which comprises biblical text, stories, and conversational data), and avenues for future improvement through data diversification and domain adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。