LLM评分器的评估结果受评判协议影响极大,需统一报告标准。
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

- 提出一套标准化的评估协议,明确评判尺度与数据处理方式
- 同一组评判结果在不同协议下准确率可从0.551跳至0.899
- 强调协议透明性,让评估结果可复现、可比较
基于评分量表的LLM作为裁判是否可替代人工标注,取决于其与人类标签的一致性。然而,相同的判断结果会因看似细微的选择——如评分尺度、保留样本、弃权与无效输出的处理方式,以及跨项目与评分维度的结论合并方式——而产生截然不同的一致性数值。这些选择背后有成熟的统计理论支持,但多源于心理测量学、计量经济学和语料标注领域,而非当前评估实践所需。本文将这些选择视为测量协议,明确报告数值所估计的内容;整合相关成果形成单一溯源分析,并应用于三项已发表的LLM裁判评估。对于非退化的二分类判断,皮尔逊相关系数、斯皮尔曼等级相关系数、肯德尔等级相关系数、φ系数与马修斯相关系数本质相同,报告多个名称实为重复同一数值。科恩κ系数仅因边际不匹配因子(0,1]差异而不同,在双方正类判定频率相等时具有相同渐近方差。在剔除样本情况下,整体准确率仅能确定为最坏情况区间,范围可达未覆盖比例。在一个含逐项人工标签的评分基准上,仅改变协议即可使报告准确率从0.551升至0.899,并令κ值跨越零点,而未更改任何一条原始判断。最终提炼出一份报告检查清单,确保一致性声明可重建、可比较。
原文摘要 · Abstract (English)
Whether a rubric-based LLM judge can replace human annotation is decided by its measured agreement with human labels. Yet the same verdicts can support wildly varying agreement numbers, depending on seemingly minor choices: the judgment scale, the retained cases, the handling of abstentions and invalid outputs, and the pooling of verdicts across items and rubric criteria. The statistics that settle these choices are established, but in psychometrics, econometrics, and corpus annotation rather than in the evaluation practice that needs them. We treat the choices as a measurement protocol that fixes what the reported number estimates before any metric is computed, assemble the relevant results into a single source-attributed analysis, and apply it to three published LLM-judge evaluations. For non-degenerate binary verdicts, Pearson's $r$, Spearman's $ρ$, Kendall's $τ_b$, the phi coefficient, and the Matthews correlation coefficient are exactly the same statistic, so reporting several repeats one number under different names. Cohen's $κ$ differs from them only through a marginal-mismatch factor in $(0,1]$ and shares their asymptotic variance when judge and human assign the positive verdict equally often. Under exclusion, accuracy over all cases is pinned down only to a worst-case interval as wide as the uncovered fraction. On a rubric benchmark carrying per-criterion human labels, protocol choice alone moves reported accuracy from $0.551$ to $0.899$ and carries $κ$ across zero, without altering a single verdict. We distill the analysis into a reporting checklist that makes agreement claims reconstructible and comparable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。