arXiv:2509.20287cs.CLcs.AI2025-09中稿 · Tenth Conference o…被引 2

研究机器翻译评估中准确与流畅的权衡,发现现有评分标准偏向准确。

Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation

  • 分析评估指标在准确与流畅间的倾向性,发现多数指标偏重准确。
  • 元评估结果显示标准测试集更青睐准确型指标,存在系统组成偏差。
  • 提出合成系统方法控制偏差,帮助更公平地比较评估指标。

我们研究了机器翻译中准确性和流畅性之间的权衡问题。研究表明,这种权衡在评估层面尤为显著,现有主流评估指标普遍倾向于准确性,其得分与翻译准确性相关性高于流畅性。更重要的是,该权衡同样存在于元评估层面,标准WMT元评估结果偏好以准确性为导向的指标而非流畅性导向的指标。这一偏差部分源于元评估数据集中所包含系统的构成。为缓解此偏差,我们提出一种在元评估中合成翻译系统的方法。研究强调理解该权衡在元评估中的重要性及其对评估指标排名的影响。

原文摘要 · Abstract (English)

We investigate the tradeoff between adequacy and fluency in machine translation. We show the severity of this tradeoff at the evaluation level and analyze where popular metrics fall within it. Essentially, current metrics generally lean toward adequacy, meaning that their scores correlate more strongly with the adequacy of translations than with fluency. More importantly, we find that this tradeoff also persists at the meta-evaluation level, and that the standard WMT meta-evaluation favors adequacy-oriented metrics over fluency-oriented ones. We show that this bias is partially attributed to the composition of the systems included in the meta-evaluation datasets. To control this bias, we propose a method that synthesizes translation systems in meta-evaluation. Our findings highlight the importance of understanding this tradeoff in meta-evaluation and its impact on metric rankings.

机器翻译评估指标元评估准确性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。