arXiv:2506.19571cs.CLcs.AI2025-06ACL被引 11

人类评估未必比自动指标更准,给翻译评测设上限。

Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress

  • 用人类标注作基准,测试自动评估指标的极限
  • 顶尖自动指标表现常与人类相当甚至更好
  • 警示:别误信评估结果,需重新思考进步衡量方式

在机器翻译评估中,自动指标的表现通常以与人工判断的一致性来衡量。近年来,自动指标与人类判断的一致性持续提升。为更清晰理解指标性能并确立上限,本文在机器翻译元评估(即对评估指标能力的评估)中引入了人类基准。结果显示,人类标注者并非始终优于自动指标,顶尖自动指标的表现常与人类基准持平或更优。尽管这些发现暗示已达人类水平,但作者仍提出多个警示。最后,探讨了研究结果对整个领域的影响:我们是否还能可靠衡量机器翻译的进步?本文旨在揭示评估能力的局限,引发对评测体系核心问题的讨论。

原文摘要 · Abstract (English)

In Machine Translation (MT) evaluation, metric performance is assessed based on agreement with human judgments. In recent years, automatic metrics have demonstrated increasingly high levels of agreement with humans. To gain a clearer understanding of metric performance and establish an upper bound, we incorporate human baselines in the MT meta-evaluation, that is, the assessment of MT metrics' capabilities. Our results show that human annotators are not consistently superior to automatic metrics, with state-of-the-art metrics often ranking on par with or higher than human baselines. Despite these findings suggesting human parity, we discuss several reasons for caution. Finally, we explore the broader implications of our results for the research field, asking: Can we still reliably measure improvements in MT evaluation? With this work, we aim to shed light on the limits of our ability to measure progress in the field, fostering discussion on an issue that we believe is crucial to the entire MT evaluation community.

机器翻译评估指标人类基准元评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。