arXiv:2510.22028cs.CL2025-10被引 3

发现翻译质量评估指标存在长度偏差,越长的译文越易被误判为差。

Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metrics

  • 通过对比多种评估方法,发现长度影响评分结果
  • 长译文即使无错也会被高估错误率,短译文更易获高分
  • 适用于数据清洗、模型优化等依赖质量评估的场景

质量评估(QE)指标在机器翻译中用于无参考评价,并日益成为数据筛选和候选重排序的依据。然而,长度偏差在这些指标中的普遍性和影响尚未得到充分研究。通过对10种不同语言对上表现优异的基于学习的和大型语言模型作为裁判(LLM-as-a-Judge)的QE指标进行系统研究,我们揭示了两个关键的长度效应:第一,无论译文质量如何,随着翻译长度增加,QE指标会持续高估错误;第二,在质量相近的候选译文中,基于学习的指标倾向于偏好较短译文,而LLM-as-a-Judge指标则从近似长度中立到偏好较长译文不等。这种依赖于指标的长度效应会导致翻译结果因长度而非质量被选中或淘汰,并可能传播至依赖QE信号的数据选择或系统优化流程中。我们追溯了MetricX-24 QE累积长度惩罚的根源,发现是训练数据中长篇无错样本分布偏斜所致。作为诊断性干预,我们在训练中引入长度归一化,该简单修改能有效解耦错误预测与序列长度的关系,使不同长度的翻译获得更可靠的评估信号。

原文摘要 · Abstract (English)

Quality Estimation (QE) metrics are vital in machine translation for reference-free evaluation and increasingly serve as selection criteria in data filtering and candidate reranking. However, the prevalence and impact of length bias in QE metrics have been underexplored. Through a systematic study of top-performing learned and LLM-as-a-Judge QE metrics across 10 diverse language pairs, we reveal two critical length effects: First, QE metrics consistently over-predict errors with increasing translation length, even for high-quality, error-free texts. Second, when candidates of comparable quality are available for the same source text, learned QE metrics tend to favor shorter translations, whereas LLM-as-a-Judge metrics range from being approximately length-neutral to favoring longer translations. These metric-dependent length effects can favor or penalize translations based on length rather than quality and can propagate into downstream pipelines that rely on QE signals for data selection or system optimization. We trace the root cause of MetricX-24 QE's cumulative-length penalty to skewed supervision distributions, in which longer error-free examples are underrepresented in training data. As a diagnostic intervention, we apply length normalization during training and show that this simple modification effectively decouples error prediction from sequence length, yielding more reliable QE signals across translations of varying length.

质量评估长度偏差机器翻译数据筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。