用大模型生成假语料,显著提升低资源语言翻译效果
Scaling Low-Resource MT via Synthetic Data Generation with LLMs
- 用大模型从欧共体语料生成跨语言合成语料,支持147对语言对
- 在7种目标语言上,合成数据使翻译性能大幅提升,即使有噪声也有效
- 开源了SynOPUS数据集,适合低资源语言研究者使用
我们研究了大语言模型生成的合成数据在低资源机器翻译中的潜力。针对七种不同目标语言,基于英语Europarl语料构建文档级合成语料,并通过回译扩展至147个额外语言对。自动与人工评估均证实其整体质量较高。我们通过(i)识别有效的训练策略,(ii)对比HPLT数据集,(iii)研究训练数据量变化的影响,以及(iv)测试其在非以英语为中心的翻译任务中的适用性,验证了其实际价值。最后,我们发布了SynOPUS——一个公开的合成平行语料库。结果表明,即使存在噪声,大模型生成的合成数据也能显著提升低资源语言的翻译性能。
原文摘要 · Abstract (English)
We investigate the potential of LLM-generated synthetic data for improving low-resource Machine Translation (MT). Focusing on seven diverse target languages, we construct a document-level synthetic corpus from English Europarl, and extend it via pivoting to 147 additional language pairs. Automatic and human evaluation confirm its overall high quality. We study its practical application by (i) identifying effective training regimes, (ii) comparing our data with the HPLT dataset, (iii) studying the effect of varying training data size, and (iiii) testing its utility beyond English-centric MT. Finally, we introduce SynOPUS, a public repository for synthetic parallel datasets. Our findings show that LLM-generated synthetic data, even when noisy, can substantially improve MT performance for low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。