LLM在动态信念追踪中表现不佳,难以回忆过去信念。
Dynamic Theory of Mind as a Temporal Memory Problem: Evidence from Large Language Models
- 将心智理论视为时间延续的记忆问题,测试模型跨轮次信念追踪能力。
- 模型能准确推断当前信念,但更新后无法回忆先前信念状态。
- 揭示了大模型在长期社交互动中的认知局限,适合关注人机交互的研究者。
心智理论(ToM)是社会认知和人机交互的核心,大型语言模型(LLMs)常被用于理解与表征ToM。然而,多数评估将ToM视为单一时刻的静态判断,主要依赖错误信念测试,忽视了其关键动态维度:对他人信念随时间演变的表征、更新与回溯能力。本文将动态ToM视为时序延展的表征记忆问题,提出DToM-Track评估框架,在受控多轮对话中检验模型对更新前信念的回忆、当前信念的推断及信念变化的检测能力。以LLMs为计算探针,发现系统性不对称:模型可可靠推断当前信念,但在信念更新后难以维持和检索先前状态。该现象在不同模型家族与规模下一致存在,符合认知科学中已知的近因偏差与干扰效应。结果表明,追踪信念轨迹远超传统错误信念推理,提示将ToM与时间表征和记忆机制关联,揭示了大模型在长期人机交互中社会推理的深层挑战。
原文摘要 · Abstract (English)
Theory of Mind (ToM) is central to social cognition and human-AI interaction, and Large Language Models (LLMs) have been used to help understand and represent ToM. However, most evaluations treat ToM as a static judgment at a single moment, primarily relying on tests of false beliefs. This overlooks a key dynamic dimension of ToM: the ability to represent, update, and retrieve others' beliefs over time. We investigate dynamic ToM as a temporally extended representational memory problem, asking whether LLMs can track belief trajectories across interactions rather than only inferring current beliefs. We introduce DToM-Track, an evaluation framework to investigate temporal belief reasoning in controlled multiturn conversations, testing the recall of beliefs held prior to an update, the inference of current beliefs, and the detection of belief change. Using LLMs as computational probes, we find a consistent asymmetry: models reliably infer an agent's current belief but struggle to maintain and retrieve prior belief states once updates occur. This pattern persists across LLM model families and scales, and is consistent with recency bias and interference effects well documented in cognitive science. These results suggest that tracking belief trajectories over time poses a distinct challenge beyond classical false-belief reasoning. By framing ToM as a problem of temporal representation and retrieval, this work connects ToM to core cognitive mechanisms of memory and interference and exposes the implications for LLM models of social reasoning in extended human-AI interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。