高质低资源英古翻译中,回译增益有限,甚至可能降低性能。
The Saturation Point of Backtranslation in High Quality Low Resource English Gujarati Machine Translation
- 用MBART50模型生成过滤后的回译数据扩充训练集
- 添加合成数据后验证集BLEU从43.8降至更低
- 揭示低资源场景下回译存在饱和现象
回译(Backtranslation, BT)在低资源机器翻译中广泛用于利用单语语料生成额外合成训练数据。尽管该方法对多数语对表现良好,但在高质量低资源场景下的有效性仍不明确。本文以多语言预训练模型MBART50为基础,研究英文到古吉拉特语的翻译任务。基线系统在约5万句对的高质量平行语料上训练,验证集BLEU达43.8。通过单语古吉拉特语文本生成并筛选回译样本进行数据增强,结果发现加入合成数据未提升性能,反而轻微下降。我们采用BLEU、ChrF++、TER、BLEURT等多指标评估,并分析饱和原因。研究表明,在特定低资源场景下,回译可能达到收益递减点,对后续研究具有启示意义。
原文摘要 · Abstract (English)
Backtranslation BT is widely used in low resource machine translation MT to generate additional synthetic training data using monolingual corpora. While this approach has shown strong improvements for many language pairs, its effectiveness in high quality, low resource settings remains unclear. In this work, we explore the effectiveness of backtranslation for English Gujarati translation using the multilingual pretrained MBART50 model. Our baseline system, trained on a high quality parallel corpus of approximately 50,000 sentence pairs, achieves a BLEU score of 43.8 on a validation set. We augment this data with carefully filtered backtranslated examples generated from monolingual Gujarati text. Surprisingly, adding this synthetic data does not improve translation performance and, in some cases, slightly reduces it. We evaluate our models using multiple metrics like BLEU, ChrF++, TER, BLEURT and analyze possible reasons for this saturation. Our findings suggest that backtranslation may reach a point of diminishing returns in certain low-resource settings and we discuss implications for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。