用回译法增强英语-卢加达语机器翻译,显著提升低资源语言表现
Data Augmentation With Back translation for Low Resource languages: A case of English and Luganda
- 通过回译生成合成双语数据,缓解低资源语言数据稀缺问题
- 在双向翻译中性能超越旧基准超10个BLEU分,多指标验证效果
- 适合低资源语言翻译研究者,方法可推广至其他小语种场景
本文研究将回译(Back translation, BT)作为半监督技术,用于提升英语-卢加达语神经机器翻译(NMT)模型的表现,重点应对低资源语言的数据短缺挑战。通过利用公开数据与网络爬取的单语语料,构建定制化NMT模型,并采用迭代式与增量式回译策略。我们创新性地在多个小型数据集上策略性选择用于增量回译的语料,形成新方法。实验结果显示,英语-卢加达语双向翻译性能均超过先前基准10个以上BLEU分数,且在SacreBLEU、ChrF2、TER等综合评估指标上表现优异,全面反映翻译质量。研究结论证实:在精心筛选数据集的前提下,回译技术能有效提升低资源语言的NMT性能,建立新基准,展现出其在该领域的巨大潜力。
原文摘要 · Abstract (English)
In this paper,we explore the application of Back translation (BT) as a semi-supervised technique to enhance Neural Machine Translation(NMT) models for the English-Luganda language pair, specifically addressing the challenges faced by low-resource languages. The purpose of our study is to demonstrate how BT can mitigate the scarcity of bilingual data by generating synthetic data from monolingual corpora. Our methodology involves developing custom NMT models using both publicly available and web-crawled data, and applying Iterative and Incremental Back translation techniques. We strategically select datasets for incremental back translation across multiple small datasets, which is a novel element of our approach. The results of our study show significant improvements, with translation performance for the English-Luganda pair exceeding previous benchmarks by more than 10 BLEU score units across all translation directions. Additionally, our evaluation incorporates comprehensive assessment metrics such as SacreBLEU, ChrF2, and TER, providing a nuanced understanding of translation quality. The conclusion drawn from our research confirms the efficacy of BT when strategically curated datasets are utilized, establishing new performance benchmarks and demonstrating the potential of BT in enhancing NMT models for low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。