arXiv:2609.05949cs.CLcs.HC2026-09

文档级翻译评估的协议存在盲点,无法识别段落间不连贯性。

The Blindness of Document-Level Translation Evaluation

论文配图:The Blindness of Document-Level Translation Evaluation
图 1 · 摘自论文原文
  • 用混拼不同系统译文的反事实测试,破坏段落一致性但保留文档呈现。
  • 一致与不一致文档的评分、排名和错误标注无统计差异。
  • 评估者能识别连贯文本,但评分结果仍无视内容逻辑断裂。

文档级机器翻译评估通过向标注者提供完整文档,假设这能引发文档级判断。我们通过反事实条件(MIX)检验该假设:将来自不同系统的段落混合组成文档,保留文档呈现形式但破坏跨段一致性。在18,420次专家英语到韩语标注及14种自动评估指标下,一致与不一致文档的得分、系统排名和错误标注在统计上无差异。感知实验显示,标注者在匹配片段中能以87.3%准确率识别出连贯文本。文档呈现确实改变了标注行为,但未反映在最终评分中。问题不在于评分偏低,而在于投入大量资源构建的文档级系统、度量和标注可能并未测量其本应衡量的内容。

原文摘要 · Abstract (English)

Document-level machine translation (MT) evaluation extends segment-level protocols by presenting full documents to annotators, on the assumption that such presentation elicits document-level judgments. We test this assumption with a counterfactual condition (MIX) in which each document combines segments drawn from different systems, preserving document-level presentation while breaking cross-segment consistency. Across 18,420 expert Englis-to-Korean annotations and 14 automatic metrics, scores, system rankings, and error annotations are statistically equivalent between coherent and incoherent documents. Perception does not explain this: shown matched passages, raters identify the coherent one as the work of a single translator in 87.3% of trials. Document presentation does change how annotators work, but that change does not reach the recorded output. What is blind is the protocol, not the annotator. The concern is not that scores fall short, but that the resources invested in document-level systems, metrics, and annotation may not be measuring what they are intended to measure.

翻译评估文档级评测盲点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。