arXiv:2509.20293cs.LGcs.AI2025-09被引 8

LLM评估基准设计缺陷导致排名看似可信实则混乱,需警惕判断噪声。

When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity

  • 通过评分框架一致性检测判断偏差,揭示评分者自相矛盾。
  • 发现多数模型排名中超过90%的评分差异无法由评分标准解释。
  • 适合关注大模型评估可靠性的研究者与评测系统设计者。

基于大语言模型(LLM)的评估基准被广泛用于衡量复杂模型行为,但其设计引入了传统基于真实标签的基准所没有的失效模式。我们指出,若缺乏明确目标与可验证构造,基准排名可能产生高度自信却实为噪声的结果。为此,我们提出两种诊断机制:结构一致性(Schematic adherence)量化评分结果中可由评分标准解释的比例,揭示评分者偏离自身规则时的未解释方差;心理测量有效性(Psychometric validity)聚合内部一致性和区分效度信号,衡量每次评测中不可消除的不确定性。应用于Arena-Hard Auto基准,我们发现主流评分者存在严重评分框架不一致和因子坍缩现象:例如,DeepSeek-R1-32B的未解释方差超过90%,多数评判维度间的因子相关性高于0.93。此外,我们还发现Arena-Hard Auto采用的ELO式聚合方法会掩盖真实的排名不确定性。结果凸显了评估设计缺陷对有效性的侵蚀,并提出了更严谨、具可靠性意识的基准构建原则。代码与数据集已开源至https://github.com/penfever/judgment-to-noise。

原文摘要 · Abstract (English)

LLM-judged benchmarks are increasingly used to evaluate complex model behaviors, yet their design introduces failure modes absent in conventional ground-truth based benchmarks. We argue that without tight objectives and verifiable constructions, benchmark rankings can produce high-confidence rankings that are in fact largely noise. We introduce two mechanisms to diagnose these issues. Schematic adherence quantifies how much of a judge's overall verdict is explained by the explicit evaluation schema, revealing unexplained variance when judges deviate from their own rubric. Psychometric validity aggregates internal consistency and discriminant validity signals to quantify irreducible uncertainty in any benchmarking run. Applying these tools to Arena-Hard Auto, we find severe schema incoherence and factor collapse across popular judges: for example, unexplained variance exceeding 90 percent for DeepSeek-R1-32B and factor correlations above 0.93 for most criteria. We also show that the ELO-style aggregation used by Arena-Hard Auto collapses and masks genuine ranking uncertainty. Our results highlight design failures that undermine validity and offer actionable principles for building better-scoped, reliability-aware LLM-judged benchmarks. We released our code and dataset at https://github.com/penfever/judgment-to-noise

大模型评估评测基准判断噪声

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。