用英语做桥梁,提升非英语间翻译质量
EnAnchored-X2X: English-Anchored Optimization for Many-to-Many Translation
- 以英语为锚点生成跨语言数据,构建双向翻译训练集
- 在72个非英语互译方向上显著提升翻译效果
- 适合需要多语种翻译能力的研究者和开发者
大语言模型在以英语为中心的语言对上表现出强大的翻译能力,但在直接的非英语(x2x)翻译中表现不佳。本文提出一种基于合成数据生成的框架,利用模型已有的英译非英语(en2x)能力,将英语平行语料扩展为全向数据集,并设计以英语为参考的质量评估代理,有效收集高质量的x2x训练数据。结合基于偏好优化的方法,该方法在72个x2x方向上显著提升主流LLM的性能,同时还能增强en2x翻译效果。结果表明,通过战略性利用英语中心优势,可有效推动大语言模型实现全面的多语言翻译能力。代码、数据集和模型检查点已开源至https://github.com/NJUNLP/EAX。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated strong machine translation capabilities for English-centric language pairs but underperform in direct non-English (x2x) translation. This work addresses this limitation through a synthetic data generation framework that leverages models' established English-to-x (en2x) capabilities. By extending English parallel corpora into omnidirectional datasets and developing an English-referenced quality evaluation proxy, we enable effective collection of high-quality x2x training data. Combined with preference-based optimization, our method achieves significant improvement across 72 x2x directions for widely used LLMs, while generalizing to enhance en2x performance. The results demonstrate that strategic exploitation of English-centric strengths can bootstrap comprehensive multilingual translation capabilities in LLMs. We release codes, datasets, and model checkpoints at https://github.com/NJUNLP/EAX
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。