arXiv:2604.01457cs.CL2026-04被引 3

发现大模型错得自信的内部机制,可精准干预校准

Wired for Overconfidence: A Mechanistic Perspective on Inflated Verbalized Confidence in LLMs

  • 从内部信号出发,定位导致自信错误的关键神经元电路
  • 中后期层中少数MLP块和注意力头主导过度自信表达
  • 干预这些模块能显著改善模型输出可信度,适合做模型调试

大型语言模型不仅常犯错,还往往表现出过度自信:在给出事实错误答案时,会以极高语气强度表达信心,而非提示不确定性。这种口头上的过度自信可能误导用户,并削弱置信度评分作为可靠不确定性指标的作用,但其内在机制仍不清楚。本文从电路级视角对大模型中的过度口头自信进行分析,围绕三个维度展开:将口头自信建模为可微分的内部信号、识别因果性地放大该信号的神经电路、并利用这些发现实现推理时的定向校准。在两个指令微调的大模型及三个数据集上,我们发现一小部分集中于中后期层的MLP块与注意力头,始终在最终词元位置写入自信膨胀信号。进一步表明,针对这些电路进行推理时干预,可显著提升模型校准效果。结果表明,大模型的口头过度自信由可识别的内部电路驱动,可通过针对性干预缓解。

原文摘要 · Abstract (English)

Large language models are often not just wrong, but \emph{confidently wrong}: when they produce factually incorrect answers, they tend to verbalize overly high confidence rather than signal uncertainty. Such verbalized overconfidence can mislead users and weaken confidence scores as a reliable uncertainty signal, yet its internal mechanisms remain poorly understood. We present a circuit-level mechanistic analysis of this inflated verbalized confidence in LLMs, organized around three axes: capturing verbalized confidence as a differentiable internal signal, identifying the circuits that causally inflate it, and leveraging these insights for targeted inference-time recalibration. Across two instruction-tuned LLMs on three datasets, we find that a compact set of MLP blocks and attention heads, concentrated in middle-to-late layers, consistently writes the confidence-inflation signal at the final token position. We further show that targeted inference-time interventions on these circuits substantially improve calibration. Together, our results suggest that verbalized overconfidence in LLMs is driven by identifiable internal circuits and can be mitigated through targeted intervention.

大模型过度自信校准机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。