arXiv:2608.23120cs.CLcs.AI2026-08

首次为英语-庞纳尔语构建统计机器翻译系统并建立基准

Statistical Machine Translation Systems of English-Pnar Language Pair : Some Insights of the Emperical Study

论文配图:Statistical Machine Translation Systems of English-Pnar Language Pair : Some Insights of the Emperical Study
图 1 · 摘自论文原文
  • 用报纸语料构建10234句平行语料,训练三组短语基SMT模型
  • 最佳模型在庞纳尔语→英语上达14.97的BLEU,首次量化基准
  • 词汇化重排提升翻译质量,低资源下MERT调优反而降低性能

庞纳尔语是印度梅加拉亚邦贾因蒂亚丘陵地区约40万人使用的南亚语系语言,缺乏数字语料与自然语言处理资源。本文首次开展英语-庞纳尔语机器翻译研究。基于Wyrta报纸文章,构建包含10,234句的平行语料库,使用Moses、GIZA++、KenLM,在三个配置下训练短语基统计机器翻译系统,每方向采用词汇化重排和最小误差率训练(MERT)调优。在371句保留测试集上评估,最优模型在庞纳尔语→英语方向达到14.97的BLEU(chrF2: 33.42,TER: 77.60),英语→庞纳尔语方向为11.16(chrF2: 31.38,TER: 93.51),确立首个定量基准。词汇化重排使庞纳尔语→英语翻译提升3.73个BLEU点,反映源语言SOV与目标语言SVO结构差异;但在低资源条件下MERT调优导致性能下降。最后分析残余错误,包括形态学词外词(OOV)、长距离重排及卡西语混用,提出未来向神经与多语言翻译发展的方向。

原文摘要 · Abstract (English)

Pnar, an Austroasiatic language spoken by approximately 0.4 million people in the Jaintia Hills of Meghalaya, lacks the digital corpora and natural language processing (NLP) resources. This paper presents the first machine translation study for the English and Pnar language pair. Using articles collected from the Wyrta newspaper, we built a parallel corpus comprising of 10,234 sentences and trained phrase-based statistical machine translation (SMT) systems the models using 9,563 parallel corpora under three configurations for each direction using Moses, GIZA++ , KenLM, varying lexicalized reordering and minimum error rate training (MERT) tuning. The models are evaluated on a held out test set of 371 sentences, the best performing system achieves a BLEU score of 14.97 (chrF2: 33.42, TER: 77.60) for Pnar to English and 11.16 (chrF2: 31.38, TER: 93.51) for English to Pnar, establishing the first quantitative benchmark for this language pair. Lexicalized reordering improves translation quality by 3.73 BLEU points for Pnar to English, reflecting the structural shift from the source language's SOV word order to the target language's SVO order, whereas MERT tuning degrades BLEU performance under low resource conditions. Finally, we analyze the remaining translation errors, including morphological out of vocabulary (OOV) words, long-distance reordering and Khasi code mixing and discuss future directions toward neural and multilingual machine translation for Pnar.

机器翻译低资源南亚语系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。