用对比判断提升机器翻译人工评估的可靠性与效率
Enhancing Human Evaluation in Machine Translation with Comparative Judgment
- 让标注员对比两个译文,比单个评分更准确
- 对比方式使标注一致性提高38.5%,错误标记更一致
- 适合需要高效、可靠评估的翻译研究者
人工评估对快速演进的语言模型至关重要,但受标注员水平和任务设计影响。本研究探索将对比判断引入机器翻译的人工标注,评估三种设置:点状多维质量度量(MQM)、并排式MQM(SxS MQM)及其简化版相对排名(SxS RR)。MQM要求标注员标记错误片段及严重程度;SxS MQM扩展为对同一输入的双译文进行成对错误标注;SxS RR仅要求选择更好译文。关键发现:(1) SxS设置的标注者间一致性高于MQM;(2) SxS MQM相比MQM在显式对比系统上平均提升38.5%、其他系统提升19.5%的跨译文错误标记一致性;(3) 所有设置均产生稳定系统排名,且SxS RR比(SxS) MQM更高效;(4) SxS设置能揭示MQM忽略的细微差异,且不改变整体系统评价。为推动研究,我们将公开包含377组中英和104组英德三重标注数据集。
原文摘要 · Abstract (English)
Human evaluation is crucial for assessing rapidly evolving language models but is influenced by annotator proficiency and task design. This study explores the integration of comparative judgment into human annotation for machine translation (MT) and evaluates three annotation setups-point-wise Multidimensional Quality Metrics (MQM), side-by-side (SxS) MQM, and its simplified version SxS relative ranking (RR). In MQM, annotators mark error spans with categories and severity levels. SxS MQM extends MQM to pairwise error annotation for two translations of the same input, while SxS RR focuses on selecting the better output without labeling errors. Key findings are: (1) the SxS settings achieve higher inter-annotator agreement than MQM; (2) SxS MQM enhances inter-translation error marking consistency compared to MQM by, on average, 38.5% for explicitly compared MT systems and 19.5% for others; (3) all annotation settings return stable system rankings, with SxS RR offering a more efficient alternative to (SxS) MQM; (4) the SxS settings highlight subtle errors overlooked in MQM without altering absolute system evaluations. To spur further research, we will release the triply annotated datasets comprising 377 ZhEn and 104 EnDe annotation examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。