大模型的自信表达更反映是否愿意提交答案,而非答案对错。
Reported Confidence in LLMs Tracks Commitment More Than Correctness

- 用两阶段弃权实验发现,模型自信度预测提交意愿比正确性更强。
- 自信度与答案正确性无关,但能准确预判模型是否愿提交回答。
- 适合关注模型决策机制与可信度评估的研究者阅读。
自信是模型对其答案正确性的概率估计。尽管口头自信报告被广泛用作大语言模型的不确定性度量,但其是否应被视为正确性估计仍不明确。我们通过神经科学中的感知决策两阶段弃权范式进行测试:模型先作答并报告自信度,再决定是否提交或放弃。在四个非推理模型、不同提示框架和自信格式下,口头自信度对提交/放弃决策的预测能力显著优于答案正确性。校准后的词元对数概率则表现出相反特征,其弃权预测与正确性判别相关,体现为答案证据信号。去除口头自信度与对数概率共享的方差后,剩余部分仍与提交意愿一致,但与正确性的关联降至接近随机水平。该分离现象在四个推理模型、四种难度不同的基准测试中均成立,涵盖困难选择题到前沿自由问答。对Gemma 3和4的机制分析显示:生成口头自信的后回答状态,在弃权提示前已编码未来弃权决策,且主要由该决策组织,而非正确性,二者在激活空间中大致正交。沿特定自信方向调控可因果改变弃权行为。因此,口头自信与对数概率不可互换:前者反映内部提交准备状态,后者追踪答案证据与正确性,挑战将口头报告视作可靠性的代理做法。
原文摘要 · Abstract (English)
Confidence is an estimate of the probability that a chosen answer is correct. Verbal confidence reports are widely used as uncertainty measures in large language models, but whether they are best understood as estimates of correctness is unclear. We test this with a two-stage abstention paradigm from the neuroscience of perceptual decision making: a model first answers and reports its confidence, then decides whether to commit it to a user or abstain. Across four non-reasoning models, prompt framings, and confidence formats, verbal confidence predicted the commit/abstain decision substantially better than whether the answer was correct. Calibrated token log-probabilities showed the opposite profile, with abstention-prediction coupled to correctness discrimination, the signature of an answer-evidence signal. After removing the variance verbal confidence shared with log-probabilities, the residual stayed aligned with commitment while its link to correctness fell to near chance. The dissociation generalised to four reasoning models across four benchmarks of varying difficulty, from hard multiple-choice to frontier-level freeform questions. Mechanistic analyses in Gemma 3 and 4 were convergent: a post-answer state known to causally support verbal-confidence generation already encoded the future abstention decision before the abstention prompt, organised mainly by that decision rather than by correctness, the two lying in approximately orthogonal directions in activation space. Steering along a verbal-confidence-specific direction causally shifted abstention. Verbal and log-probability confidence are thus not interchangeable: log-probabilities track answer evidence and correctness, whereas verbal confidence is better understood as a behaviour-facing readout of an internal commit-readiness state, challenging the practice of treating verbal reports as proxies for reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。