arXiv:2604.17393cs.CL2026-04ACL被引 2

发现翻译评估指标在新领域表现不佳,人类标注者意见也不一致。

Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains

论文配图:Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains
图 1 · 摘自论文原文
  • 构建跨领域错误片段标注数据集,固定标注者控制变量
  • 指标与人类在新闻域一致性达0.69,但新领域仅0.78-0.83
  • 建议用指标-人类一致率对比标注者间一致率来评估

自动评估指标在机器翻译发展中的作用至关重要,但其在领域迁移下的鲁棒性尚不明确。现有指标多基于WMT基准,引发对未见领域适应能力的担忧。以往研究常混杂系统、标注者或评估条件差异,难以分离领域影响。为此,我们构建了多标注者跨领域错误片段标注数据集(CD-ESA),包含18.8k条人工错误片段标注,覆盖三个语对,在一个已见新闻域和两个未见技术域上评估六套翻译系统。结果表明,自动指标在段落级对齐上看似稳健(最高0.69一致性),但一旦考虑标注者变异,其稳健性大幅下降。平均标注可提升标注者间一致性达+0.11。在未见化学领域,指标表现显著弱于人类(标注者间一致性0.78–0.83,人类达0.96)。建议评估跨领域表现时,应比较指标-人类一致率与标注者间一致率,而非仅看原始指标-人类一致性。

原文摘要 · Abstract (English)

Automatic evaluation metrics are central to the development of machine translation systems, yet their robustness under domain shift remains unclear. Most metrics are developed on the Workshop on Machine Translation (WMT) benchmarks, raising concerns about their robustness to unseen domains. Prior studies that analyze unseen domains vary translation systems, annotators, or evaluation conditions, confounding domain effects with human annotation noise. To address these biases, we introduce a systematic multi-annotator Cross-Domain Error-Span-Annotation dataset (CD-ESA), comprising 18.8k human error span annotations across three language pairs, where we fix annotators within each language pair and evaluate translations of the same six translation systems across one seen news domain and two unseen technical domains. Using this dataset, we first find that automatic metrics appear surprisingly robust to domain-shifts at the segment level (up to 0.69 agreement), but this robustness largely disappears once we account for human label variation. Averaging annotations increases inter-annotator agreement by up to +0.11. Metrics struggle on the unseen chemical domain compared to humans (inter-annotator agreement of 0.78-0.83 vs. 0.96). We recommend comparing metric-human agreement against inter-annotator agreement, rather than comparing raw metric-human agreement alone, when evaluating across different domains.

机器翻译评估指标领域泛化人类标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。