只在关键少数问题上验证,能极大提升LLM评测效率
Blind to the Pivotal Vote: Aggregate Independence Metrics Miss Where Verification Actually Helps

- 聚焦仅一票之差的难题,验证信号才真正有效
- 关键问题上准确率提升10.4至23.3个百分点
- 适合想减少验证开销的研究者和开发者
LLM评委小组是标准评估工具,但先前研究发现评委错误高度相关:九名评委的有效信息仅相当于两名独立评委,聚合也仅小幅缩小差距。引入新证据源(如执行测试套件)在大规模下对有效投票数无显著影响(-0.04,95%置信区间[-0.10, +0.02])。聚合依赖性与条件决策效用是不同问题。基础多数决规则表明,仅一票之差的决策可被替换。实证显示,准确率提升完全集中于这些关键查询,幅度达+10.4至+23.3个百分点,其他地方无增益。该模式在三个代码基准和四种评委规模(9人扩展及56次依赖采样检查)中均成立,增益为+6.5至+16.1个百分点。在HumanEval+/MBPP+上,多数侧替换规则将整体准确率从82.44%提升至85.62%,仅需在16.2%的查询上调用信号;而仅使用信号的策略更优,达87.60%。因此,群体层面依赖诊断与边际分层效用互补,受影响集特征可生成任意单票替换策略的调用缩减规则。
原文摘要 · Abstract (English)
LLM judge panels are a standard evaluation tool, but prior work reports highly correlated panel errors: nine judges provide roughly the effective information of two independent ones, and aggregation closes only a small fraction of the gap. A natural remedy--a signal from a different evidence source, e.g., executing a test suite--produced no distinguishable change in the panel's effective-vote count at scale (-0.04, 95\% CI [-0.10, +0.02]). Aggregate dependence and conditional decision utility are different questions. Elementary majority arithmetic fixes the affected set for single-ballot substitution: only decisions with a one-vote margin can change. The empirical question is whether panel error rates rise and useful substitutions concentrate there. They do: the entire accuracy gain concentrates on these pivotal queries, where it is large (+10.4 to +23.3 percentage points across three headline configurations), and is exactly zero elsewhere. We confirm the pattern across three code benchmarks and four panel sizes (a 9-judge extension and 56 dependent subsampling checks, gain +6.5 to +16.1 percentage points). On HumanEval+/MBPP+, a majority-side replacement rule raises overall accuracy from 82.44\% to 85.62\% while invoking the signal on 16.2\% of queries; signal-only remains stronger at 87.60\%. Thus population-level dependence diagnostics and margin-stratified utility are complementary, and the affected-set characterization yields a call-reduction rule for any specified single-ballot substitution policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。