用神经科学框架揭示大模型自信值背后的计算机制。
The Computational Basis of Confidence in Large Language Models

- 以答案对数差为信号,检验其是否反映潜在决策变量。
- 多任务验证中,对数差呈现符合规范的自信模式。
- 适用于非推理类多模态模型,为可信部署提供依据。
可靠的自信度——即模型判断自身答案正确的概率——是语言模型可信部署的关键。现有研究主要评估自信度对正确性的预测能力与校准性,但未回答一个更根本的问题:自信信号本身代表什么?答案对数可能反映可计算最优自信的潜在决策变量,也可能仅是非贝叶斯组合证据的启发式偏好信号。本文采用计算神经科学中的统计决策自信(SDC)框架,将答案对数差(LD)视为潜在决策变量的候选读出信号,检验了SDC预测的定性特征。在三个感知判别任务和一个记忆决策任务中,涵盖三种多模态非推理模型与一种推理模型,LD均满足这些特征,包括诊断性正确/错误折叠X型模式,表明在此类场景下,答案对数作为潜在决策变量的单调读出信号而非启发式偏好。在复杂视觉推理任务中,LD仍能超越客观难度预测正确性,但完整几何特征缺失,说明当前框架在缺乏显式规范过程模型时存在边界。研究为多模态语言模型的自信提供了计算解释,明确了答案对数作为决策变量读出的适用条件,并确立SDC作为生物与人工智能中研究自信的统一框架。
原文摘要 · Abstract (English)
Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and whether it is calibrated, leaving open a more fundamental question: what does the confidence signal itself represent? Answer logits may reflect a latent decision variable sufficient to compute normative confidence, or instead a heuristic preference signal that combines the available evidence in a non-Bayesian manner. We address this using statistical decision confidence (SDC), a normative framework from computational neuroscience. Treating the answer-logit difference (LD) as a candidate readout of the latent decision variable, we test the qualitative signatures predicted by SDC. Across three perceptual discrimination tasks and a memory-based decision task, spanning three multimodal non-reasoning models and one reasoning model, LD satisfied these signatures -- including the diagnostic correct/error folded-X pattern -- showing that, in these settings, answer logits behave as monotonic readouts of a latent decision variable rather than heuristic preference scores. In complex visual reasoning, LD continued to predict correctness beyond objective task difficulty, but the full geometric signatures of SDC were absent, illustrating the current boundary of the framework when explicit normative process models are unavailable. These results provide a computational account of confidence in multimodal language models, delineate when answer logits behave as readouts of a latent decision variable, and establish SDC as a unifying framework for studying confidence across biological and artificial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。