用语法错误数据提升低资源语言翻译效果,让模型更懂目标语言。
NSL-MT: Linguistically Informed Negative Samples for Efficient Machine Translation in Low-Resource Languages
- 生成目标语言的语法错误样本,强制模型拒绝不合理输出。
- 1000条数据训练效果相当于普通方法5000条,性能提升3%-89%。
- 适合平行语料稀缺但需高精度翻译的场景。
我们提出负空间学习机器翻译(NSL-MT),一种针对低资源语言的训练方法。该方法通过合成目标语言的语法违规样本,增强有限的平行语料,并在模型对这些语言学无效输出赋予高概率时进行显式惩罚。NSL-MT在所有测试基线中均取得改进,对表现良好的模型带来3%-12%的BLEU提升,对初始支持不足的模型则实现56%-89%的显著增益。此外,NSL-MT实现了5倍的数据效率提升:仅用1000个训练样本即可达到常规训练5000个样本的效果。因此,该方法为平行语料稀缺场景提供了高效可靠的训练替代方案。
原文摘要 · Abstract (English)
We introduce negative space learning machine translation (NSL-MT), a training method for underresourced languages, that augments limited parallel data with synthetically generated violations of the target language's grammar and explicitly penalizes the model when it assigns high probability to these linguistically invalid outputs. NSL-MT delivers improvements across all baselines we tested, including 3-12% BLEU gains for well-performing models and 56-89% gains for models lacking decent initial support. Furthermore, NSL-MT provides a 5x data efficiency multiplier: training with 1,000 examples matches or exceeds normal training with 5,000 examples. NSL-MT thus provides a data-efficient alternative training method for settings where parallel data is limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。