arXiv:2603.22161cs.LG2026-03被引 4

发现大模型用自信度决定是否回答,而非仅依赖输出概率。

Causal Evidence that Language Models use Confidence to Drive Behavior

  • 通过四阶段实验验证模型用内部自信信号控制回答行为。
  • 提升或抑制自信信号可直接改变拒绝回答的比例。
  • 适合关注模型自主决策与可信推理的研究者。

元认知——评估自身认知表现的能力——在多个物种中引导适应性行为。已有大量研究证明可从语言模型输出中提取自信信号,但核心问题仍存:模型是否真的利用这些信号来控制行为,例如决定是否回答或放弃?为此,我们设计了四阶段范式。第一阶段在无放弃选项下获取基线自信估计。第二阶段显示,大模型在决定放弃时采用隐式自信阈值,其效果量比其他机制高一个数量级。第三阶段通过激活操控提供直接因果证据:增强或抑制自信信号分别导致放弃率下降或上升。第四阶段系统性调节指令阈值,证明大模型能主动运用自信信号实施放弃策略。关键发现是,除了基于校准对数概率的自信外,言语自信也能独立预测放弃行为,尽管其客观上对答案正确性区分能力较弱。在最后一句前标记处的激活解码显示,这两种可观测指标都是更丰富内部表征的有损读出。结果表明,放弃行为不仅由输出分布中的证据强度决定,更由多维内部自信表征与阈值策略共同作用解释——支持大模型具备结构化元认知控制,这在模型向自主代理演进的背景下日益重要。

原文摘要 · Abstract (English)

Metacognition -- assessing the quality of one's own cognitive performance -- guides adaptive behavior across species. Substantial research demonstrates that confidence signals can be extracted from language model outputs, yet a fundamental question remains: do models actually use these signals to control behavior, such as deciding whether to answer or abstain? To investigate, we developed a four-phase paradigm. Phase~1 elicited baseline confidence estimates without an abstention option. Phase~2 revealed that LLMs apply an implicit threshold to internal confidence when deciding to abstain, with confidence effect sizes approximately an order of magnitude larger than alternative mechanisms. Phase~3 provided direct causal evidence through activation steering: boosting or suppressing confidence signals correspondingly decreased or increased abstention rates. Phase~4 extended this by systematically varying instructed thresholds, demonstrating that LLMs actively deploy confidence signals to implement abstention policies. Critically, beyond calibrated log-probability based confidence derived from the output distribution, verbal confidence independently predicted abstention across all models, despite being objectively less discriminatory of answer correctness. Activation decoding at the last pre-answer token further showed that both observable measures are lossy readouts of a richer internal representation. Together, these results suggest that abstention is not fully captured by the strength of evidence in the output distribution alone, but is better explained by the joint operation of a multidimensional internal confidence representation and threshold-based policies -- consistent with structured metacognitive control in LLMs, a capacity of growing importance as models transition to autonomous agents that must recognize their own uncertainty.

大模型元认知行为控制自信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。