arXiv:2602.12150cs.AIcs.CL2026-02被引 1

GPT-4o虽能模仿人类社交判断,但缺乏稳定的心理理论机制。

GPT-4o Lacks Core Features of Theory of Mind

  • 基于认知科学定义,测试模型对心理状态如何影响行为的因果理解。
  • 在逻辑等价任务中表现失败,行为预测与心理推断一致性低。
  • 揭示当前大模型社交能力依赖表面模式,非真正心理建模能力。

大型语言模型(LLMs)是否具备心理理论(ToM)?现有研究多通过基准测试评估其在社交任务中的表现,发现模型在多项任务中表现良好。然而,这些评估并未检验ToM所主张的核心表征:即心理状态与行为之间的因果模型。本文基于认知基础的ToM定义,提出并测试了一种新评估框架,旨在探究模型是否具备一致、通用且连贯的心理状态导致行为的因果模型——无论该模型是否符合人类心理。结果表明,尽管模型在简单心理理论范式中能近似人类判断,但在逻辑等价任务中表现失败,且其行为预测与对应心理状态推断之间的一致性较低。因此,这些发现表明,当前大模型展现出的社交能力并非源于通用或一致的心理理论。

原文摘要 · Abstract (English)

Do Large Language Models (LLMs) possess a Theory of Mind (ToM)? Research into this question has focused on evaluating LLMs against benchmarks and found success across a range of social tasks. However, these evaluations do not test for the actual representations posited by ToM: namely, a causal model of mental states and behavior. Here, we use a cognitively-grounded definition of ToM to develop and test a new evaluation framework. Specifically, our approach probes whether LLMs have a coherent, domain-general, and consistent model of how mental states cause behavior -- regardless of whether that model matches a human-like ToM. We find that even though LLMs succeed in approximating human judgments in a simple ToM paradigm, they fail at a logically equivalent task and exhibit low consistency between their action predictions and corresponding mental state inferences. As such, these findings suggest that the social proficiency exhibited by LLMs is not the result of a domain-general or consistent ToM.

心理理论大模型评测认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。