pairwise 比较能准确反映生成模型真实性能,效果优于直接评分。
Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings

- 用配对比较和Elo算法评估生成模型性能
- Elo排名与真实准确率相关性超0.9,显著高于直接评价
- 风格和评委偏见影响小,回声重复是偏好主因
配对比较结合Elo等聚合方法已成为生成模型评估的核心手段,但人们担忧其可能偏向表面风格特征或受评委偏见影响。本文发现,当有真实准确率作为基准时,配对比较所得模型排名与真实排名高度一致。我们将五个知名基准转化为自由形式生成评估,结果显示Elo排名与准确率排名的斯皮尔曼相关系数超过0.9,且在评委能力较弱时显著优于直接评估。此外,尽管多数判断发生在两个答案均正确(或均错误)的配对中,风格和评委偏见对排名影响甚微。我们进一步发现,答案末尾的重复(回声)是导致评委偏好的因果因素。
原文摘要 · Abstract (English)
Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge biases. In a more positive turn, we show that model rankings from pairwise comparisons strongly agree with ground-truth-based accuracy rankings when such ground truth is available for comparison. By converting five well-known benchmarks into free-form generative evaluations, we find that Elo rankings achieve a Spearman correlation above 0.9 with accuracy rankings and substantially outperform direct evaluation when the judge is weak. Furthermore, style and judge bias have only minor effects on model rankings, despite most judgments occurring on pairs where both candidate answers are correct (or incorrect). On such pairs, we find that repetition after the final answer (echo) is a causal driver of judge preference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。