arXiv:2605.09826cs.AIcs.MA2026-05被引 2

测试智能体在真实场景中理解他人隐含认知的能力

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents

论文配图:EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents
图 1 · 摘自论文原文
  • 构建300个动态任务,模拟有遮挡、私密信息和通信限制的3D家庭环境
  • 前沿模型在功能型任务上全数失败,平均准确率仅0.0%,而字面信念题达45.0%
  • 揭示93%失败源于认知协作断裂,为未来研究提供明确改进方向

心智理论(ToM)使人类能高效协作。在多智能体场景中,AI代理同样需要此能力,但现有基准大多仅测试字面式ToM,通过直接提问信念。在具身环境中基于隐含信念做出最优行动的能力,即功能型ToM,仍缺乏有效评估。我们提出EnactToM,一个包含300个具身多智能体任务的动态基准,设置于3D家庭环境,具有部分可观测性、私密信息和通信约束。每个任务均经形式化验证以确保可解性和所需认知深度,新任务会随模型性能提升而生成以增加难度。在困难划分上,所有七种评估的前沿模型在功能任务完成上的通过率均为0.0%,而字面信念探测平均为45.0%。人工分析显示,93%的采样失败源于认知协作断裂,如信息隐瞒、忽略伙伴约束和消息误分配,为未来工作提供了具体目标。

原文摘要 · Abstract (English)

Theory of Mind (ToM), the ability to track others epistemic state, makes humans efficient collaborators. AI agents need the same capacity in multi agent settings, yet existing benchmarks mostly test literal ToM by asking direct belief questions. The ability act optimally on implicit beliefs in embodied environments, called functional ToM, remains largely untested. We introduce EnactToM, an evolving benchmark of 300 embodied multi-agent tasks set in a 3D household with partial observability, private information, and constrained communication. Each task is formally verified for solvability and required epistemic depth, and new tasks are generated increase difficulty as models improve. On the hard split, all seven evaluated frontier models score 0.0% Pass^3 on functional task completion, while averaging 45.0% on literal belief probes. Manual analysis traces 93% of sampled failures to epistemic coordination breakdowns such as withheld information, ignored partner constraints, and misallocated messages, providing a concrete target for future work.

心智理论具身智能多智能体评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。