arXiv:2608.20975cs.AI2026-08

测试视觉语言模型能否把心理理论转化为协同社交行为。

Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models

论文配图:Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models
图 1 · 摘自论文原文
  • 设计多模态互动基准MOSAIC,融合语言、动作、眼神与表情。
  • 13个模型在200次试验中均无法按心理预期产生协调行为。
  • 显式加入心理推理模块的PCM-LLM表现最佳,说明信念-行动耦合关键。

有效社交互动要求智能体将心理状态推断转化为言语与非言语渠道的协同行为信号。然而现有评估体系将心理理论(ToM)推理与具身行为分开测评,未能衡量推理到行动之间的差距。本文提出MOSAIC(多模态社交行动、推理与沟通编排)基准,让两个具身智能体在合作与竞争场景中交互,需整合语言陈述、空间轨迹、注视方向和面部表情,并在系统性变化的ToM约束下完成任务。对13个模型(含11个视觉语言模型,VLMs)进行每模型200次试验后发现,多数VLM无法生成符合预期结果的行为,且强制设定的ToM层级约束并未带来可靠行为响应。信号层面分析揭示双重瓶颈:多数模型无法生成方向一致的非言语信号;即使存在信号,也难以理解他人行为并作出反应。作为结构参考的PCM-LLM(带显式心理模块)在所有条件下均成功,表明显式的信念-行为耦合是此类任务的关键要素。

原文摘要 · Abstract (English)

Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channels simultaneously. Yet existing benchmarks evaluate theory of mind (ToM) reasoning and embodied behavior in isolation, leaving unmeasured the gap between social inference and social action. We introduce MOSAIC (Multimodal Orchestration of Social Action, Inference, and Communication), a controlled benchmark in which two embodied agents interact across cooperative and competitive scenarios requiring integration of verbal statements, spatial trajectories, gaze direction, and facial expression under systematically varied ToM constraints. Evaluating 13 models, including 11 VLMs, across 200 trials per model, we find that VLMs fail to produce behaviors consistent with the expected outcomes under ToM-order constraints, and that imposing explicit ToM-order constraints produces no reliable behavioral change aligned with the specified reasoning level. Signal-level analysis reveals two sequential bottlenecks: most models cannot produce directionally coherent nonverbal signals, and even when signals are present, VLM agents fail to interpret others behaviors and react to them. PCM-LLM, included as a structured architectural reference point with an explicit ToM module, succeeds across all conditions, suggesting that explicit belief-action coupling is a sufficient ingredient for this class of tasks.

心理理论多模态交互视觉语言模型社会智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。