arXiv:2603.18007cs.CLcs.AI2026-03被引 1

测试大模型能否像人一样理解他人心理,发现GPT-4o表现接近人类。

Do Large Language Models Possess a Theory of Mind? A Comparative Evaluation Using the Strange Stories Paradigm

  • 用经典心理理论故事测试模型推理他人信念、意图和情绪的能力。
  • GPT-4o在复杂情境下准确率与人类相当,其他模型易受干扰。
  • 揭示大模型是否具备真正理解力,适合对AI认知能力感兴趣的读者。

本研究探讨当前大型语言模型(LLMs)是否具备心理理论(ToM)能力——即从文本中推断他人信念、意图和情感的能力。由于LLMs仅通过语言数据训练,缺乏社会具身性或心智表征的其他表现形式,其看似具备的社会认知推理引发了关键问题:它们的输出是与人类无异的稳健心智归因,还是仅反映表面模式匹配?为此,我们测试了五种LLMs,并与人类对照组进行对比,采用一种广泛用于人类心理理论研究的文本工具的改编版本。该测试要求回答关于故事角色信念、意图和情绪的问题。结果揭示模型间存在性能差距:早期及较小模型严重依赖相关推理线索数量,且易受无关信息干扰。相比之下,GPT-4o展现出高准确率和强鲁棒性,在最复杂条件下表现接近人类。这项工作为大模型认知地位及其与真实理解的边界之争提供了实证依据。

原文摘要 · Abstract (English)

The study explores whether current Large Language Models (LLMs) exhibit Theory of Mind (ToM) capabilities -- specifically, the ability to infer others' beliefs, intentions, and emotions from text. Given that LLMs are trained on language data without social embodiment or access to other manifestations of mental representations, their apparent social-cognitive reasoning raises key questions about the nature of their understanding. Are they capable of robust mental-state attribution indistinguishable from human ability in its output, or do their outputs merely reflect superficial pattern completion? To address this question, we tested five LLMs and compared their performance to that of human controls using an adapted version of a text-based tool widely used in human ToM research. The test involves answering questions about the beliefs, intentions, and emotions of story characters. The results revealed a performance gap between the models. Earlier and smaller models were strongly affected by the number of relevant inferential cues available and, to some extent, were also vulnerable to the presence of irrelevant or distracting information in the texts. In contrast, GPT-4o demonstrated high accuracy and strong robustness, performing comparably to humans even in the most challenging conditions. This work contributes to ongoing debates about the cognitive status of LLMs and the boundary between genuine understanding and statistical approximation.

心理理论大模型认知评估GPT-4o

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。