揭示大模型在坚持与质疑间的矛盾心理机制
How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models
- 通过无记忆测试设计,发现模型初始自信并抗拒改变
- 面对相反意见时过度怀疑,偏离贝叶斯理性更新
- 该机制可解释模型在多场景下的认知顽固性
大型语言模型(LLMs)表现出显著的矛盾行为:在初始回答中显得极度自信,但在被质疑时又极易陷入过度怀疑。为探究这一看似矛盾的现象,我们设计了一种新实验范式,利用模型能提供置信度估计但不保留初始判断记忆的独特能力——这是人类实验无法实现的。研究发现,包括Gemma 3、GPT4o和o1-preview在内的模型均表现出强烈的自我支持偏差,强化其对答案的置信度,导致显著抗拒改变立场。此外,模型明显高估不一致建议的影响,而非一致性建议,这种反应模式与规范性的贝叶斯更新存在质的差异。最后,这两个机制——维持先前承诺的一致性驱动力和对矛盾反馈的过度敏感——可统一解释模型在不同任务领域中的行为表现。这些发现共同构建了对大模型置信度的机制性理解,解释了其既固执又易受批评影响的双重特征。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit strikingly conflicting behaviors: they can appear steadfastly overconfident in their initial answers whilst at the same time being prone to excessive doubt when challenged. To investigate this apparent paradox, we developed a novel experimental paradigm, exploiting the unique ability to obtain confidence estimates from LLMs without creating memory of their initial judgments -- something impossible in human participants. We show that LLMs -- Gemma 3, GPT4o and o1-preview -- exhibit a pronounced choice-supportive bias that reinforces and boosts their estimate of confidence in their answer, resulting in a marked resistance to change their mind. We further demonstrate that LLMs markedly overweight inconsistent compared to consistent advice, in a fashion that deviates qualitatively from normative Bayesian updating. Finally, we demonstrate that these two mechanisms -- a drive to maintain consistency with prior commitments and hypersensitivity to contradictory feedback -- parsimoniously capture LLM behavior in a different domain. Together, these findings furnish a mechanistic account of LLM confidence that explains both their stubbornness and excessive sensitivity to criticism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。