评估大模型判官时,仅看一致性不够,还得测它对关键变化的敏感度。
A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation
- 用不变性S和敏感性R双维度衡量大模型判官的评价有效性
- 高一致性的判官在构造改变时仅31.9%会改判,说明敏感度不足
- 建议同时报告S与R,且检查标签数据是否被表面特征误导
LLM-as-a-judge评估通常依赖一致性和对表面扰动的鲁棒性,但这些无法保证构念效度。我们提出将评估者构念效度定义为二维指标:不变性S(在保持构念的修改下判别结果不变的概率)和敏感性R(在最小构念改变下判别结果变化的概率)。实验表明S与R相互独立,任何单一数值都无法完整反映比较关系。我们在7个判官、4个领域中,使用7种构念改变干预和5种仅注册控制,由人工标注确定干预方向,并将生成、验证与评判分配给不同模型家族。在匹配的不变性水平S ≥ 0.90下,平均不变性S = 0.945,但敏感性R = 0.319。敏感性在范围与强度修改间存在差异:R_scope = 0.383 vs R_strength = 0.262,差距+0.121,所有7个判官均同向。进一步审计5个公开标签集发现,仅基于表面特征的预测器可复现配对模式下55%-67%的标签,包括MT-Bench人类投票的67.4%。结果表明,高一致性可与弱构念敏感性共存,亟需联合报告不变性与敏感性,并审计验证集本身。
原文摘要 · Abstract (English)
LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but reliability does not establish construct validity. We formalize construct validity for an evaluator as a two-dimensional profile: invariance S, the probability that a verdict is unchanged under construct-preserving edits, and construct sensitivity R, the probability that it changes under minimal construct-changing edits. We show that S and R are independent and that no scalar summary preserves all relevant comparisons. We measure the profile across 7 judges and 4 domains using 7 construct-changing intervention types and 5 register-only controls, with intervention direction determined by human annotators and generation, verification, and judging assigned to disjoint model families. At matched invariance S >= 0.90, judges average S = 0.945 but R = 0.319. Sensitivity also differs between scope and strength edits: R_scope = 0.383 versus R_strength = 0.262, a +0.121 gap with the same sign for all 7 judges. We further audit five public label sets and find that surface-only predictors reproduce 55%-67% of labels in paired mode, including 67.4% of MT-Bench human votes. These results show that high judge agreement can coexist with weak sensitivity to changes in the construct being evaluated, motivating joint reporting of invariance and sensitivity and auditing the validation set itself.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。