arXiv:2409.09598cs.CLcs.AI2024-09中稿 · WMT 2024被引 29

提出新评估方法SPA,让自动评分更接近人类判断且结果更可信。

Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy

  • 基于配对准确率改进,融合人类与评分的统计显著性。
  • 相比旧方法,SPA能区分更多不同评分模型,减少误判。
  • 已在2024年WMT评测中正式使用,适合语言生成评估研究者。

选择最贴近人工评判的自动评估指标往往困难,因缺乏明确的“最佳模拟”定义。需借助元指标比较人工判断与自动评分,但指标排序依赖元指标选择。本文提出软配对准确率(Soft Pairwise Accuracy, SPA),在配对准确率基础上引入人类判断与自动评分的统计显著性。实验表明,SPA对评估系统/片段数量变化更稳定;传统配对准确率仅能产生有限的离散输出值,导致多个指标被错误赋予相同分数;而SPA有效解决了该问题,具备更强区分力,能实现更显著的指标间比较。SPA已被选为2024年WMT指标共享任务的官方系统级评估指标。

原文摘要 · Abstract (English)

Selecting an automatic metric that best emulates human annotators is often non-trivial, because there is no clear definition of "best emulates." A meta-metric is required to compare the human judgments to the automatic metric scores, and metric rankings depend on the choice of meta-metric. We propose Soft Pairwise Accuracy (SPA), a new meta-metric that builds on Pairwise Accuracy (PA) but incorporates the statistical significance of both the human judgments and the metric scores. We show that SPA is more stable than PA with respect to changes in the number of systems/segments used for evaluation. We also show that PA can only assign a small set of distinct output values to metrics, and this results in many metrics being artificially assigned the exact same PA score. We demonstrate that SPA fixes this issue. Finally, we show that SPA is more discriminative than PA, producing more statistically significant comparisons between metrics. SPA was selected as the official system-level metric for the 2024 WMT Metrics Shared Task.

评估方法自动评分统计显著

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。