拆解大模型自洽中错误一致性的来源,发现其主要由机械性偏好决定。
Decomposing Wrong-Consensus Agreement in LLM Self-Consistency
- 通过分解一致性成分,分离出仅由答案偏好驱动的机械部分
- 开放权重模型中机械成分覆盖率达99%以上,而前沿模型残留显著
- 揭示了自洽性并非可靠证据,适合关注模型可信度的读者
语言模型重复采样结果的一致性常被当作答案可靠的依据,但错误答案也能表现出同样强的一致性。本文提出量化分解方法,将多元一致性指数Gamma(归一化参考尺度d=(1-p)/(C-1))分解为机械成分(仅由单个问题的答案偏好决定)与未解释残差。机械参考不泄露信息:每个样本的偏好与准确率仅基于其其他运行估计。在公开的GPT-4.1每轮数据上,多选题GPQA-Diamond的覆盖率phi为0.81–0.93,开放域AIME为0.59–0.78,残差达1.54–2.80 Gamma单位,超过校准后的运行级偏好异质性参考。在五组开放权重检查点(Qwen3.5-9B/122B、Qwen3.8-27B、Gemma4-26B/31B)上,固定协议(每题4次运行,K=32票)下的控制复现显示所有十组细胞机械覆盖率接近100%(phi≈1,小幅度超调与有限捐赠插值偏差一致),且对两轮设计稳健;最大细胞(qwen3.5-122b, p=0.222)仍达到饱和(phi=1.041),位于GPT-4.1 AIME准确率区间内。跨系统对比显示,在相近总体准确率下,开放权重模型近似完全机械性一致,而前沿家族模型残差更大。该差异因采样协议设计而混淆。一致性是分级证据,而非认证。未提出新投票方法;代码与证据已公开。
原文摘要 · Abstract (English)
Agreement among repeated samples of a language model is routinely read as evidence about answer reliability, yet wrong answers can agree just as strongly as right ones. This paper asks what information wrong-consensus agreement actually contains, and answers with a quantitative decomposition. A pluralistic agreement index Gamma, normalized by the reference scale d=(1-p)/(C-1), is split into a mechanical component (agreement delivered by a per-case answer preference alone) and a preference-unexplained residual. The mechanical reference is leak-free: each case's preference and accuracy are estimated from its other runs only. On public GPT-4.1 per-run data, coverage phi (the mechanical/empirical ratio) shows a benchmark-associated direction: 0.81-0.93 on multiple-choice GPQA-Diamond against 0.59-0.78 on open-domain AIME, where a residual of 1.54-2.80 Gamma units survives, more than absorbed by a calibrated run-level preference-heterogeneity reference. A controlled replication under one fixed protocol (four runs per question, K=32 votes) on five open-weights checkpoints (Qwen3.5-9B/122B, Qwen3.8-27B, Gemma4-26B/31B) finds near-complete mechanical coverage in all ten cells (phi approximately 1, with a small overshoot consistent with a quantified finite-donor plug-in bias), robust to a two-run design; the largest cell (qwen3.5-122b, p=0.222) sits inside the GPT-4.1 AIME accuracy range and still saturates (phi=1.041). A cross-system contrast at comparable aggregate accuracy contrasts near-complete mechanical agreement in the open-weights models against a larger preference-unexplained residual in the frontier family. This contrast is confounded with sampling protocol by design. Agreement is graded evidence, not certification. No new voting method is proposed; code and evidence are committed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。