arXiv:2608.18106cs.CLcs.AI2026-08

揭示大模型过度自信的根源,发现其倾向于默认输出确定性答案。

Different Facets of Verbalised Overconfidence: an Interpretability Study

论文配图:Different Facets of Verbalised Overconfidence: an Interpretability Study
图 1 · 摘自论文原文
  • 通过控制推理场景,分析模型在三种不确定表达方式下的行为差异。
  • 数值置信度评分下过自信现象最严重,不确定性仅由少数专用特征实现。
  • 干预特定特征可缓解过自信问题,且效果跨语言和任务通用。

大型语言模型常表现出过度自信,即使证据表明应持保留或回避态度。本文以Qwen3-4B为研究对象,在控制逻辑必然性与可能性的推理场景下,考察其通过言语认知标记、回避回答和数值置信度分数三种方式表达不确定性的行为。结果证实模型存在明显过度自信倾向,尤其在要求输出数值置信度时更为显著。我们提出一种可解释性方法,能差异化识别负责不确定性和确定性的转码特征。分析显示,Qwen3-4B默认机制通过一组广泛共享的特征生成确定性输出,而不确定性则作为稀疏覆盖,由少量专用特征驱动。干预这些不确定性特征不仅因果验证了过自信的内在失衡,还能有效缓解错误。该特征集在三种表达方式、多种语言及分布外模态任务中均具泛化能力。

原文摘要 · Abstract (English)

Large language models tend to overconfidence, giving assertive answers when the evidence suggests hedging or abstention. Using controlled reasoning scenarios that manipulate logical necessity and possibility, we study this behavior in Qwen3-4B, across three ways to express uncertainty: verbal epistemic markers, abstention, and numeric confidence scores. Our results confirm this tendency toward overconfidence, particularly when the model is prompted to output a numeric confidence score. At the interpretability level, we propose a method that differentially identifies transcoder features responsible for uncertainty and certainty. Our analysis reveals Qwen3-4B's default mechanism favors certainty generation through a broad coalition of shared features, while uncertainty is implemented as a sparse override mediated by a small set of dedicated features. Intervening on these uncertainty features both causally proves this imbalance underlying overconfidence and also mitigate overconfident errors. The same set of features generalise across the three uncertainty-expression settings, languages, and an out-of-distribution modality task.

大模型过自信可解释性不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。