评测智能体在复杂社交场景中的多视角心智推理能力
SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
- 构建动态多智能体环境生成多模态社交数据,支持第一/第三人称双视角评估
- 人类在双视角任务中准确率远超大模型,差距达40.1%(第一人称)和26.4%(第三人称)
- 适合研究具身智能、社会推理与视觉语言模型的局限性
人类在动态真实社交互动中持续通过感知环境推断他人状态、目标与行为。然而,现有心智理论(ToM)评测多局限于静态文本场景,与真实交互存在显著差距。本文提出SoMi-ToM评测基准,用于评估具身多智能体复杂社交互动中的多视角心智推理能力。该基准基于交互环境SoMi生成的丰富多模态数据,涵盖多样化的制作目标与社会关系。评估框架支持多层级:(1)第一人称评估提供任务中的多模态输入(视觉、对话、动作等),实现实时状态推断;(2)第三人称评估提供任务后完整的视频与文本记录,用于目标与行为推断。此方法兼顾主观即时体验与客观全局观察,全面检验模型能力。我们构建了包含35个第三人称视频、363张第一人称图像及1225个专家标注的多选题(每题三选项)的数据集。在该数据集上,系统评估了人类与多个前沿视觉-语言大模型(LVLMs)的表现。结果显示,LVLM在两项评估中均显著落后于人类:第一人称平均准确率差距为40.1%,第三人称为26.4%。这表明未来视觉-语言模型需在具身复杂社交互动中进一步提升心智推理能力。
原文摘要 · Abstract (English)
Humans continuously infer the states, goals, and behaviors of others by perceiving their surroundings in dynamic, real-world social interactions. However, most Theory of Mind (ToM) benchmarks only evaluate static, text-based scenarios, which have a significant gap compared to real interactions. We propose the SoMi-ToM benchmark, designed to evaluate multi-perspective ToM in embodied multi-agent complex social interactions. This benchmark is based on rich multimodal interaction data generated by the interaction environment SoMi, covering diverse crafting goals and social relationships. Our framework supports multi-level evaluation: (1) first-person evaluation provides multimodal (visual, dialogue, action, etc.) input from a first-person perspective during a task for real-time state inference, (2) third-person evaluation provides complete third-person perspective video and text records after a task for goal and behavior inference. This evaluation method allows for a more comprehensive examination of a model's ToM capabilities from both the subjective immediate experience and the objective global observation. We constructed a challenging dataset containing 35 third-person perspective videos, 363 first-person perspective images, and 1225 expert-annotated multiple-choice questions (three options). On this dataset, we systematically evaluated the performance of human subjects and several state-of-the-art large vision-language models (LVLMs). The results show that LVLMs perform significantly worse than humans on SoMi-ToM: the average accuracy gap between humans and models is 40.1% in first-person evaluation and 26.4% in third-person evaluation. This indicates that future LVLMs need to further improve their ToM capabilities in embodied, complex social interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。