测试大模型是否具备稳定的核心信念,发现它们仍无法像人一样坚持基本认知。
Do LLMs have core beliefs?

- 用对抗性对话树框架测试五大学科领域的信念稳定性
- 多数模型在持续对话中最终放弃关键事实,显示世界观不稳固
- 虽近年模型辩论能力提升,但缺乏人类特有的信念坚守机制
大型语言模型(LLMs)的兴起引发了关于其是否具备人类认知水平的讨论。然而,当前研究较少关注人类认知中一个结构性要素——核心信念:那些构成世界观基础、不易被推翻的真理。一旦放弃这些信念,将意味着对现实理解的根本转变。本文通过名为对抗性对话树(ADTs)的探测框架,在科学、历史、地理、生物和数学五个领域考察了大模型是否拥有类似核心承诺。结果表明,大多数模型无法维持稳定的认知体系。尽管部分近期模型表现出更好的稳定性,但在持续对话压力下仍会动摇关键信念。这说明尽管模型在论辩能力上有所进步,但所有现有模型仍缺失人类认知的关键特征。
原文摘要 · Abstract (English)
The rise of Large Language Models (LLMs) has sparked debate about whether these systems exhibit human-level cognition. In this debate, little attention has been paid to a structural component of human cognition: core beliefs, truths that provide a foundation around which we can build a worldview. These commitments usually resist debunking, as abandoning them would represent a fundamental shift in how we see reality. In this paper, we ask whether LLMs hold anything akin to core commitments. Using a probing framework we call Adversarial Dialogue Trees (ADTs) over five domains (science, history, geography, biology, and mathematics), we find that most LLMs fail to maintain a stable worldview. Though some recent models showed improved stability, they still eventually failed to maintain key commitments under conversational pressure. These results document an improvement in argumentative skills across model generations but indicate that all current models lack a key component of human-level cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。