arXiv:2606.31916cs.CL2026-06被引 1

测试大模型通过行动影响他人信念的能力,发现其在非对话场景下已有初步社会推理能力。

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action

论文配图:Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
图 1 · 摘自论文原文
  • 设计新框架让模型通过移动物体或引导角色达成信念目标
  • GPT-5在80%任务中成功,优于人类平均水平
  • 模型更擅长诱导真实信念,提示对齐方向积极

大型语言模型(LLMs)的理论心理(ToM)评估通常依赖被动问答,但随着模型向更自主智能体发展,亟需新评测方式。本文提出非对话规划理论心理(NCP-ToM),评估模型通过行动而非对话影响其他智能体信念状态的能力。我们构建了新框架NCP-ExploreToM,将传统任务改为:给定一组信念目标,要求模型通过移动物体或引导角色进入房间来实现目标。在600个任务实例上评估了六款前沿模型(包括GPT-5、Gemini 2.5 Pro和Claude 4系列)及人类对照组。结果显示,GPT-5在代理设置中成功完成约80%的任务,是唯一表现超过人类的模型,但在跨情境鲁棒性上仍弱于人类。所有模型与人类一样,在诱导真实信念任务上表现优于虚假信念任务,这是对齐努力的积极信号。研究揭示了大模型在非对话任务中日益显现的社会推理能力,并强调必须采用代理式评估以理解自主社会智能体的安全性与对齐性。

原文摘要 · Abstract (English)

Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in increasingly agentic and autonomous forms demands new evaluations. In this paper we evaluate an agent's ability to induce specific belief states in other agents by taking actions rather than using conversational persuasion, a capability we call Non-Conversational Planning ToM (NCP-ToM). NCP-ToM is likely to be essential for many agent use-cases, including within user-assistant interactions and pedagogical contexts, but may also present manipulation or misinformation risks. Using a novel framework, NCP-ExploreToM, we subvert the conventional task structure by providing models with a set of belief state goals and requiring them to move objects or direct characters into rooms to achieve their goals. We evaluated six frontier models, including GPT-5, Gemini 2.5 Pro and the Claude 4 series, and a cohort of human participants, across 600 task instances. GPT-5 was successful on approximately 80% of tasks in the agentic setting, and was the only model to outperform human participants on our task, but was still less robust than humans across contexts. We additionally found that all models, like humans, performed better on tasks inducing true belief states than false belief states, which is a positive signal for alignment efforts. These findings highlight emerging social-reasoning capabilities in LLMs for non-conversational task completion and underscore the necessity of agentic evaluations for understanding the safety and alignment of autonomous social agents.

理论心理智能体信念诱导安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。