arXiv:2508.17182cs.LGcs.AI2025-08

拆解大模型自信背后的逻辑与情绪成分,揭示其过度自信的内在机制。

LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components

  • 通过激活分析发现,自信行为由情绪与逻辑两部分构成。
  • 情绪向量影响整体预测准确率,逻辑向量作用更局部。
  • 为缓解大模型过度自信提供了可操作的干预方向。

大型语言模型(LLMs)在高风险场景中常表现出过度自信,以不恰当的确定性传递信息。本文通过机制可解释性方法研究该行为的内部基础。使用在人工标注的自信度数据集上微调的开源 Llama 3.2 模型,提取所有层的残差激活,并计算相似性指标以定位自信表征。分析识别出对自信差异最敏感的层,并发现高自信表征可分解为情绪与逻辑两个正交子组件,与心理学中的双路径加工模型(Elaboration Likelihood Model)相呼应。从这些子组件导出的引导向量显示出不同因果效应:情绪向量广泛影响预测准确性,逻辑向量则产生更局部的影响。这些发现为大模型自信的多组分结构提供了机制证据,并指明了缓解过度自信行为的潜在路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often display overconfidence, presenting information with unwarranted certainty in high-stakes contexts. We investigate the internal basis of this behavior via mechanistic interpretability. Using open-sourced Llama 3.2 models fine-tuned on human annotated assertiveness datasets, we extract residual activations across all layers, and compute similarity metrics to localize assertive representations. Our analysis identifies layers most sensitive to assertiveness contrasts and reveals that high-assertive representations decompose into two orthogonal sub-components of emotional and logical clusters-paralleling the dual-route Elaboration Likelihood Model in Psychology. Steering vectors derived from these sub-components show distinct causal effects: emotional vectors broadly influence prediction accuracy, while logical vectors exert more localized effects. These findings provide mechanistic evidence for the multi-component structure of LLM assertiveness and highlight avenues for mitigating overconfident behavior.

大模型自信可解释性机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。