arXiv:2608.20925cs.CLcs.AI2026-08

现有翻译评估方法忽视源文本,导致评价不公。

Source-Free MT Evaluation Is Not MT Evaluation

论文配图:Source-Free MT Evaluation Is Not MT Evaluation
图 1 · 摘自论文原文
  • 主张以源文本为基准评估翻译准确度
  • 指出当前混合指标过度依赖参考译文
  • 呼吁将质量估计作为核心评估方法

基于参考译文的评测指标仍是机器翻译评估的标准,部分原因在于质量估计方法与人工判断的相关性较低。因此,即使缺乏参考译文,源文本无关的参考译文评估已成为常态,但这违背了翻译准确性的定义,且对保持源义但与参考不同的系统不公平。本文认为,准确性必须基于源文本判断:参考译文只是源文本的一种可能表达,可能引入偏差、不充分或错误。只有当评估者将参考视为辅助证据而非主要标准时,源-参考-假设评估才公平。否则,即便考虑源文本,仍会将准确性简化为对参考译文的偏好。我们发现现有混合指标高度依赖参考译文而非源文本。这并非否定所有自动评测指标使用源文本,而是强调:任何移除源文本或允许参考主导源文本的评估协议,在结构上都无法完整衡量准确性。然而,现有论文普遍偏好参考译文指标,仅在无参考时才使用质量估计。因此,我们呼吁将质量估计重新定位为以源文本为基础的准确性评估的首要方法,而非因缺少参考而退而求其次的备选方案。同时,应设计明确优先考虑源-假设忠实度的混合指标,并将参考译文仅作为补充证据。

原文摘要 · Abstract (English)

Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less well with human judgments. As a result, source-free, reference-based evaluation has become the practical norm, even though it is unfaithful to the definition of translation adequacy and unfair to systems whose outputs preserve the source meaning while differing from the reference. This paper argues that adequacy must be judged with respect to the source. A reference is only one possible rendering of the source and may introduce bias, under-specification, or errors. We further argue that source-reference-hypothesis evaluation is fair only when the judge treats the reference as auxiliary evidence rather than as the primary standard. Otherwise, even source-aware evaluation can reduce adequacy to preference towards reference. We show the existing hybrid metrics are highly reliant on reference compared to source. Our argument is not that all automatic MT metrics fail to use the source. Rather, we argue that any evaluation protocol that removes the source, or allows the reference to dominate the source, is structurally incomplete for adequacy evaluation. However, existing MT papers generally prefer reference-based metrics and use QE metrics only when reference is unavailable. We therefore call for QE to be reframed as a primary approach to source-grounded adequacy evaluation, rather than as a fallback motivated by missing references. We further call for hybrid metrics whose designs explicitly prioritize source--hypothesis faithfulness while using references only as complementary evidence.

机器翻译评估方法质量估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。