arXiv:2605.29800cs.CL2026-05被引 14

九个大模型评估组实则仅相当于两个独立判断,因错误高度重合导致效果大打折扣。

Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels

论文配图:Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
图 1 · 摘自论文原文
  • 用有效样本量和康多塞模型量化评估组信息价值
  • 9个模型仅贡献约2个独立判断,准确率低8-22个百分点
  • 错误相关性是核心瓶颈,扩容或换算法也难补足差距

LLM作为裁判的评估小组通过聚合多个模型投票来提升可靠性,但其有效性依赖于模型间的独立性。我们构建框架测量此类小组的真实信息价值,并量化其与理想独立投票之间的差距。在三个自然语言推理数据集(每项100个真人标注)上测试由7个模型家族的9个前沿大模型组成的小组,发现其实际提供信息量仅相当于约2个独立投票。约四分之三的名义独立性因模型在相同样本上犯相同错误而丧失。结果显著:小组实际准确率比独立投票预期低8至22个百分点,且最佳单个模型在所有条件下均不低于或优于整个小组。无论增加评委数量还是使用更优聚合算法,均无法改善——现有方法最多仅能缩小11%的差距,即便已知正确答案亦如此。我们采用Kish有效样本量(n_eff)与康多塞零模型验证结论,发现该缺陷在不同提示、温度、思维链推理及配对偏好任务(RewardBench)下均稳定存在。根本瓶颈在于评委间错误相关性,而非聚合算法,表明扩大评估组无法替代真正独立的评价。

原文摘要 · Abstract (English)

LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a framework to measure the true informational value of such panels and quantify how far their reliability falls short of the independent-voting ideal. Testing a panel of 9 frontier LLMs from 7 model families on three natural language inference datasets (each with 100 human annotations per item), we find that the 9 judges effectively provide only about 2 independent votes' worth of information. Roughly three-quarters of the panel's nominal independence is lost because the models make the same mistakes on the same items. The consequences are stark: the panel's actual accuracy falls 8-22 percentage points short of what independent voting would achieve, and the best single judge matches or outperforms the full panel across all conditions. Neither adding more judges nor using smarter aggregation algorithms helps -- established methods close at most 11% of this gap, even with access to the correct answers. We quantify these findings using the Kish effective sample size (n_eff) and a Condorcet null model, and show the deficit is robust across prompt variants, temperatures, chain-of-thought reasoning, and a pairwise preference task (RewardBench). The bottleneck is correlated judges, not the aggregation algorithm, implying that scaling up panels cannot substitute for genuinely independent evaluation.

大模型评估评估偏差错误相关性模型独立性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。