小模型在复杂科学问题上,多数投票反而降低准确率,信心门也失效。
When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs

- 用多数投票提升小模型准确率会适得其反
- 即使高置信度样本,答案仍可能与多数矛盾
- 适合关注推理可靠性与模型评估的读者
对小型指令微调模型而言,基于多数投票的自一致性方法在大多数GPQA钻石级问题上降低了准确率:Qwen2.5-7B为56.6%,Llama-3-8B为65.7%。看似合理的无验证器置信度门也失败了。分析揭示三种信号失效机制,其中基于分词熵的门因测量对象错误而无效——平均约602个词元的统计量反映的是叙述流畅性而非答案置信度。在Qwen2.5-7B-Instruct-Turbo上,64次采样下198个问题中,答案与问题多数相悖的样本仍以中位数20.52纳特的置信度输出,75.7%高于10纳特。这些数据为预注册并单次测试,均通过检验。整体上,该置信度差值使正确与错误样本分离达+0.0604([+0.0183, +0.1017]),但逐题配对分析未显著区分,结果跨零点。核心观点并非概率无信息,而是跨题有效信号对题内决策几乎无用。多数一致门失败机制仍未阐明,列为开放问题。本研究仅基于单一模型,另一次注册复现因无法评估被放弃。此外,三款原生推理模型在低预算无服务器推理环境下均无法运行,原因各异,但均可下载,说明按令牌计费接口的可用性受限于实际计算边界。
原文摘要 · Abstract (English)
Self-consistency via majority vote reduces per-problem accuracy on most GPQA Diamond problems for small instruction-tuned models: 56.6% of problems for Qwen2.5-7B and 65.7% for Llama-3-8B. The obvious remedy is a verifier-free confidence gate. This version reports that the most natural repair also fails, and separates three signal failures that v1 treated as one. A token-entropy gate fails for a measurement reason: averaged over a chain of some 602 tokens, the statistic is a measurement of the prose rather than of confidence in the answer. On Qwen2.5-7B-Instruct-Turbo, 198 problems at 64 samples each, a sample whose answer contradicts its own problem's plurality still emits that answer at a median margin of 20.52 nats, with 75.7% above 10 nats. Both quantities were pre-registered and tested once on 69 problems no exploratory analysis had read; both passed. The unit is the whole result: pooled across the benchmark the margin separates correct from incorrect samples by +0.0604 on the fraction above 10 nats [+0.0183, +0.1017], excluding zero; per-problem and paired it does not, at -0.0168 [-0.0527, +0.0182], crossing zero. The claim is not that token log-probabilities carry no information, but that a signal with real across-question discrimination is close to useless for the within-question decision a router faces. The plurality-agreement gate's failure remains without a mechanism, and we report it as an open problem. These new claims rest on one model: a registered second-model replication was sampled and could not be evaluated, and we report that rejection rather than the result. We separately report that on hosted serverless inference at a small budget, three reasoning-native models could not be evaluated, for three separately measured reasons; all three are downloadable, so this bounds what a metered per-token API buys rather than what is knowable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。