arXiv:2606.22633cs.AI2026-06

研究大模型如何处理与自身观点冲突的信息,发现其内部不确定性可预测说服程度。

Confident but Conflicted: Internal Uncertainty and Cognitive Dissonance Resolution in LLMs

论文配图:Confident but Conflicted: Internal Uncertainty and Cognitive Dissonance Resolution in LLMs
图 1 · 摘自论文原文
  • 通过调节信息权威性和证据质量,测试模型在12类健康议题上的认知冲突化解行为。
  • 提出信任弹性(TE)指标,发现错误观点下所有模型的说服力趋近于零。
  • 内部不确定性指标与模型行为变化相关,为改进模型提供新方向。

大型语言模型(LLMs)常遭遇与其先前输出相矛盾的输入,如用户质疑、检索文档或网络搜索结果。尽管此类冲突的解决过程——我们称之为认知失调化解——已有行为层面的描述,但其与模型内部不确定性的关联尚不明确。为此,我们在12个具有不同知识可信度层级的健康科学命题上,沿两个维度(信息源权威性与证据质量)系统性地施加说服干预。冲突可被化解、引发反效果或导致免疫。我们引入信任弹性(Trust Elasticity, TE),一个受经济学启发的度量指标,用于衡量模型被相反证据说服的难易程度。在四个主流大模型中,TE存在显著差异;而对明显错误的命题,所有模型的TE均接近零。在两个开源模型上进一步发现,该差异与两类互补的内部不确定性指标相关:Qwen的置信度校准偏差和Llama的内部不确定性变化。这些结果将模型间的行为差异与可测量的内部属性关联起来,提示未来可通过调控内部不确定性来优化模型表现。

原文摘要 · Abstract (English)

Large language models (LLMs) frequently encounter inputs that disagree with their prior outputs, through user pushback, retrieved documents, or web search results. While the way they resolve such conflicts -- a process we frame as cognitive dissonance resolution -- has been characterized behaviorally, its connection to internal model uncertainty is not well understood. To study this systematically, we vary persuasion attempts along two dimensions, source authority and evidence quality, across 12 health-science claims of stratified epistemic status. Dissonance can be resolved through persuasion, backfire, or immunity. We introduce Trust Elasticity (TE), an econometrics-inspired measure of how readily a model is persuaded toward conflicting evidence. Across four LLMs, TE varies substantially, while clearly false claims elicit near-zero TE across all models. On two open-weight models, we further find that this variation is associated with two complementary internal uncertainty indicators, Confidence Miscalibration in Qwen and Internal Uncertainty Change in Llama. These results link cross-model behavioral variation to a measurable internal property and point to interventions targeting internal uncertainty as future work.

大模型认知失调不确定性信任弹性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。