arXiv:2503.24013cs.CL2025-03中稿 · COLM被引 5

翻译评估不能只看一个分数,准确性和自然性存在本质权衡。

You Cannot Feed Two Birds with One Score: the Accuracy-Naturalness Tradeoff in Translation

  • 用信息论证明准确与自然存在不可调和的权衡
  • 实证发现优化BLEU会先提升自然度,过度优化则严重降低自然度
  • 建议改用二维平面评估,而非单一数值比较

翻译的目标是将源语言文本转化为目标语言中既忠实于原意又表达自然的文本。然而,机器翻译研究常使用单一评分同时衡量语义准确性和表达自然度。本文基于信息论,数学证明并实证展示了这种单一评分无法全面反映系统真实表现:准确与自然之间存在固有权衡。通过评估WMT24共享任务的提交结果,我们验证了该现象。这解释了常见观察——针对特定准确率指标(如BLEU)优化系统,初期可提升自然度,但过度拟合反而显著损害自然度。因此,我们主张改变评估方式:不应仅以单一数值比较系统,而应在准确-自然度平面上进行多维对比。

原文摘要 · Abstract (English)

The goal of translation, be it by human or by machine, is, given some text in a source language, to produce text in a target language that simultaneously 1) preserves the meaning of the source text and 2) achieves natural expression in the target language. However, researchers in the machine translation community usually assess translations using a single score intended to capture semantic accuracy and the naturalness of the output simultaneously. In this paper, we build on recent advances in information theory to mathematically prove and empirically demonstrate that such single-score summaries do not and cannot give the complete picture of a system's true performance. Concretely, we prove that a tradeoff exists between accuracy and naturalness and demonstrate it by evaluating the submissions to the WMT24 shared task. Our findings help explain well-known empirical phenomena, such as the observation that optimizing translation systems for a specific accuracy metric (like BLEU) initially improves the system's naturalness, while ``overfitting'' the system to the metric can significantly degrade its naturalness. Thus, we advocate for a change in how translations are evaluated: rather than comparing systems using a single number, they should be compared on an accuracy-naturalness plane.

机器翻译评估方法权衡分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。