大模型心智推理能力随训练逐步发展,但易受语言干扰影响。
Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models

- 从发育视角追踪模型在不同训练阶段的心理状态推理能力变化。
- 模型需足够规模与训练量,且在后训练阶段才显著提升虚假信念判断能力。
- 情境建模先于心智推理出现,但对反事实信息仍敏感,暴露其脆弱性。
近期研究指出大型语言模型(LLMs)能感知文本描述中代理的信念状态,通过假信念任务(FBT)衡量,但其构念效度仍存疑。本文采用**发育视角**,追踪Olmo2与Pythia模型系列在多个训练阶段的心理状态推理行为及可能前提条件。结果发现,超越随机水平的FBT表现依赖于模型规模与充分训练量,于预训练后期才逐渐出现,并在最能诊断心智化能力的“假信念-隐含”条件下,经后训练干预(SFT、DPO)改善最明显。然而,FBT表现具有脆弱性:使用非真值动词(如“认为”)会即使在真信念条件下也增加假信念归因。为理解这些发现,我们追踪了**情境建模**的涌现——即报告描述场景基本事实属性的能力。情境建模准确率通常早于且高于FBT准确率,但情境表征在某些方面仍表现出惊人不连贯:当被问及始终知晓物品真实位置的对立者(Antagonist)的知识状态时,Olmo2 13b模型始终受到目标者(Target)知识状态和非真值动词的影响。综合来看,更大、充分训练的模型以发育适宜顺序构建部分连贯的情境模型,但表现出意外脆弱性,凸显发育与压力测试方法在评估大模型能力中的价值。
原文摘要 · Abstract (English)
Recent work suggests that Large Language Models (LLMs) are sensitive to the belief states of agents described by text, as measured by the false belief task (FBT), yet persistent concerns of construct validity remain. We adopt a **developmental perspective**, tracing the pattern of mental state reasoning behavior -- and likely **preconditions** for this behavior -- across multiple training stages in the Olmo2 and Pythia language model suites. We find that above-chance FBT performance depends both on model size and sufficient training volume, emerges relatively late in pretraining, and is most improved by post-training interventions (SFT, DPO) in the condition most diagnostic of mentalizing (False Belief, Implicit). However, FBT performance is fragile: consistent with past work, the use of non-factive verbs (e.g., thinks) increases false belief attributions even in the True Belief condition. To contextualize these findings, we track the emergence of **situation modeling**: the ability to report on basic factual properties of a described scene. Situation modeling accuracy generally precedes and exceeds FBT accuracy, yet situational representations also prove surprisingly incoherent in certain respects: when asked about the knowledge states of the Antagonist agent -- who always knows the item's true location -- Olmo2 13b is consistently influenced both by the Target agent's knowledge state and the presence of non-factive verbs. Together, these results suggest that larger, sufficiently trained models build partially coherent situation models in a developmentally appropriate sequence, yet display surprising fragility -- highlighting the value of developmental and stress-testing approaches for evaluating LLM capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。