arXiv:2603.25052cs.CLcs.AI2026-03被引 6

让大模型的自信表达更真实,解决自说自话的偏差问题

Closing the Confidence-Faithfulness Gap in Large Language Models

  • 发现自信表达与实际准确率在模型内部正交,互不干扰
  • 推理过程会破坏自信信号,导致自信心失真
  • 提出两阶段调节方法,显著提升模型自信表达的准确性

大型语言模型的自信表达常与其真实准确率脱节,但其内在机制仍不清晰。本文通过线性探针和对比激活添加(CAA)操控,发现校准能力与自信信号在模型中呈线性编码且正交,这一现象在三个开源模型和四个数据集上均成立。有趣的是,当模型同时进行推理并输出自信评分时,推理过程会干扰自信方向,加剧校准偏差,我们称之为「推理污染效应」。基于此,提出一种两阶段自适应调节管道,读取模型内部的准确率估计并修正输出,显著提升各模型的校准一致性。

原文摘要 · Abstract (English)

Large language models (LLMs) tend to verbalize confidence scores that are largely detached from their actual accuracy, yet the geometric relationship governing this behavior remain poorly understood. In this work, we present a mechanistic interpretability analysis of verbalized confidence, using linear probes and contrastive activation addition (CAA) steering to show that calibration and verbalized confidence signals are encoded linearly but are orthogonal to one another -- a finding consistent across three open-weight models and four datasets. Interestingly, when models are prompted to simultaneously reason through a problem and verbalize a confidence score, the reasoning process disrupts the verbalized confidence direction, exacerbating miscalibration. We term this the "Reasoning Contamination Effect." Leveraging this insight, we introduce a two-stage adaptive steering pipeline that reads the model's internal accuracy estimate and steers verbalized output to match it, substantially improving calibration alignment across all evaluated models.

大模型可信度可解释性校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。