arXiv:2602.22072cs.CLcs.AI2026-02

测试大模型在假信念任务中的心理推理能力,发现扰动下性能骤降。

Understanding Artificial Theory of Mind: Perturbed Tasks and Reasoning in Large Language Models

  • 通过设计扰动任务评估大模型心理建模能力
  • 所有模型在扰动下思维能力显著下降
  • 思维链提示能提升但需谨慎使用,避免误判

心智理论(ToM)指代理对他人内在状态的建模能力。为探讨大语言模型是否具备真正的ToM能力,本研究通过在假信念任务中引入扰动,考察其鲁棒性,并检验思维链提示(CoT)对性能提升及决策解释的作用。我们构建了一个手工标注、内容丰富的ToM数据集,包含经典与扰动后的假信念任务,对应正确任务完成的合理推理路径空间、后续推理忠实度、任务解法,并提出评估推理链正确性及最终答案对推理轨迹忠实度的指标。结果表明,所有评估的大型语言模型在任务扰动下均出现显著性能下降,质疑了其存在稳健的ToM能力。尽管CoT提示总体上以可解释方式提升了性能,但对某些扰动类别反而降低准确率,说明需选择性应用。

原文摘要 · Abstract (English)

Theory of Mind (ToM) refers to an agent's ability to model the internal states of others. Contributing to the debate whether large language models (LLMs) exhibit genuine ToM capabilities, our study investigates their ToM robustness using perturbations on false-belief tasks and examines the potential of Chain-of-Thought prompting (CoT) to enhance performance and explain the LLM's decision. We introduce a handcrafted, richly annotated ToM dataset, including classic and perturbed false belief tasks, the corresponding spaces of valid reasoning chains for correct task completion, subsequent reasoning faithfulness, task solutions, and propose metrics to evaluate reasoning chain correctness and to what extent final answers are faithful to reasoning traces of the generated CoT. We show a steep drop in ToM capabilities under task perturbation for all evaluated LLMs, questioning the notion of any robust form of ToM being present. While CoT prompting improves the ToM performance overall in a faithful manner, it surprisingly degrades accuracy for some perturbation classes, indicating that selective application is necessary.

心智理论大模型推理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。