让标注员同时对比多个翻译,提升评估效率和一致性
Contrastive ESA: Human Evaluation of Multiple Translations at Once
- 并列展示多个翻译版本,通过对比判断错误位置
- 标注时间减少,噪声降低,评分更稳定
- 适合需要可靠模型排序的翻译评估场景
当前机器翻译的人工评估通常单独评价单个输出,存在标注噪声高、成本大的问题。本文提出对比式错误片段标注(cESA):将同一源文本的多个翻译结果并列呈现,标注员识别主要和次要错误区域,并在0%至100%的绝对尺度上打分。通过共享多输出上下文,cESA促进更一致、高效的判断。我们在12个模型的英译日语大规模人工评估中验证该方法,结果显示相比传统逐项评估,标注耗时与噪声显著下降。不同于已有对比排序方法,cESA生成绝对质量评分,可直接进行简单、可解释的非参数模型排序,无需后期校正。
原文摘要 · Abstract (English)
Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protocol that presents multiple translations of the source input (text, video, audio, image). In cESA, the annotator sees multiple translations of the same document, marks major and minor error spans, and then assigns a score from 0% to 100% on absolute scale. By allowing annotators to access the shared context across multiple outputs, cESA facilitates more consistent and efficient judgments. We validate cESA using a large-scale human evaluation of English->Japanese translations of 12 models, demonstrating reductions in annotation time and noise compared to standard pointwise evaluation. Unlike existing contrastive ranking methods, cESA yields absolute quality judgments that enable simple, interpretable non-parametric model rankings without the need for post-hoc corrections.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。