用传话游戏生成数据,提升机器翻译评估效果
Evaluating Language Translation Models by Playing Telephone
- 通过多轮语言转换自动生成评估数据
- 在长文本和文学翻译任务上优于xCOMET
- 适合需要高质量评估的翻译研究者
当前机器翻译模型的性能已远超评估能力,制约了其在长篇及文学翻译等复杂任务上的进一步优化。我们提出一种无监督方法,通过源语言与目标语言间多次循环翻译,生成适用于不同文档长度和应用领域的训练数据。利用模型旋转和语言翻译两种方式生成的机械文本,训练评估系统,在两项任务上表现更优:(i)评分给定翻译与人工参考之间的质量差异;(ii)判断两个翻译中哪个更接近原始源文档。实验表明该方法在多个任务上超越了主流评估系统xCOMET。
原文摘要 · Abstract (English)
Our ability to efficiently and accurately evaluate the quality of machine translation systems has been outrun by the effectiveness of current language models--which limits the potential for further improving these models on more challenging tasks like long-form and literary translation. We propose an unsupervised method to generate training data for translation evaluation over different document lengths and application domains by repeated rounds of translation between source and target languages. We evaluate evaluation systems trained on texts mechanically generated using both model rotation and language translation approaches, demonstrating improved performance over a popular translation evaluation system (xCOMET) on two different tasks: (i) scoring the quality of a given translation against a human reference and (ii) selecting which of two translations is generationally closer to an original source document.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。