arXiv:2606.19544cs.CL2026-06被引 6

大规模测试发现大模型裁判存在评估偏差,需改用更科学的验证方法。

Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias

论文配图:Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias
图 1 · 摘自论文原文
  • 采用三种评估协议,系统检验21个大模型裁判的可靠性与偏见。
  • 精确匹配率虚高33至41个百分点,真实一致性远低于表面表现。
  • 适合关注模型评估可信度的研究者和开发者参考。

LLM-as-a-Judge已成为语言模型评估的主流范式,但实践中依赖的精确匹配一致率未校正随机因素,系统性夸大模型判别能力。本文开展迄今最大规模的系统性评估:涵盖来自九家供应商的21个裁判模型,在MT-Bench、JudgeBench和RewardBench三个基准上,通过三类协议(一致性、稳定性、偏见审计)进行118次运行,共约54.1万次独立判断。四项发现贯穿全样本,包括2026年4月前沿结果:精确匹配与Cohen's kappa之间存在普遍性κ值压缩(在MT-Bench上达33–41个百分点),裁判排名在不同基准间最多变动14位;两个上线部署裁判虽具高重测信度(>0.95),却存在严重位置偏见(>0.10),形成一致性-偏见悖论;在单一两两评分标准下,冗长性偏见极小(<0.011)。据此提出最小可行验证协议。

原文摘要 · Abstract (English)

LLM-as-a-Judge has become the dominant evaluation paradigm for language models, but judge validation in practice relies on exact-match agreement, a metric that does not correct for chance and systematically overstates discriminative ability. We present the largest systematic evaluation of LLM-as-a-Judge to date: 21 judges from nine providers across MT-Bench, JudgeBench, and RewardBench, evaluated under three protocols (agreement, consistency, bias audit) over 118 runs and approximately 541,000 individual judgments. Four findings emerge, consistent across the full cohort, including the April 2026 frontier: kappa deflation between exact match and Cohen's kappa is universal (33--41 pp on MT-Bench), judge rankings shift by up to 14 positions across benchmarks, high test--retest reliability (>0.95) coexists with severe position bias (>0.10) in two production-deployed judges (instantiating a consistency--bias paradox), and verbosity bias is small (<0.011) across our cohort under a single pairwise rubric. We distill these into a Minimum Viable Validation Protocol.

模型评估大模型裁判一致性偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。