arXiv:2509.25546cs.CL2025-09EMNLP被引 1

用段落差异相关性评估机器翻译,更准更稳。

Don't Sweat the Small Stuff: Segment-Level Meta-Evaluation Based on Pairwise Difference Correlation

  • 基于段落间差异而非原始分数计算相关性。
  • 在WMT'24上正确排序主流评估指标,贴近人工判断。
  • 抗噪声和系统偏差,对异常值敏感,适合质量分析。

本文提出一种新型段落级元评估指标Pairwise Difference Pearson(PDP),用于机器翻译评估,解决了以往基于Pearson ρ和Kendall τ的元评估方法的局限性。PDP采用成对差异而非原始得分,利用所有段落信息获得更稳健的评分分布理解,并通过段落内成对差异优化全局Pearson相关性。在WMT'24共享任务中的分析表明,PDP能正确排序主流评估指标,且与人工错误权重更一致。噪声注入实验显示,PDP对随机噪声、段落偏倚和系统偏倚具有鲁棒性,同时对极端异常值保持敏感。

原文摘要 · Abstract (English)

This paper introduces Pairwise Difference Pearson (PDP), a novel segment-level meta-evaluation metric for Machine Translation (MT) that address limitations in previous Pearson's $ρ$-based and and Kendall's $τ$-based meta-evaluation approaches. PDP is a correlation-based metric that utilizes pairwise differences rather than raw scores. It draws on information from all segments for a more robust understanding of score distributions and uses segment-wise pairwise differences to refine Global Pearson to intra-segment score comparisons. Analysis on the WMT'24 shared task shows PDP properly ranks sentinel evaluation metrics and better aligns with human error weightings than previous work. Noise injection analysis demonstrates PDP's robustness to random noise, segment bias, and system bias while highlighting its sensitivity to extreme outliers.

机器翻译元评估相关性段落级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。