让视觉语言机器人具备揣测他人和自我的心智能力,做出更智能的决策。
MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents
- 构建以机器人为中心的心智推理框架,同步建模自我与他人心理状态。
- 在决策和动作生成上优于GPT-4o,分别提升12.77%和12.49%。
- 适用于需要理解人类意图的机器人交互场景,如服务、协作任务。
心智理论(ToM)指推断他人信念、欲望和意图等心理状态的能力。当前视觉语言具身智能体缺乏基于心智理论的决策能力,且现有基准仅关注人类心理状态,忽视智能体自身视角,导致行为生成不连贯。为此,我们提出MindPower——一个融合感知、心智推理、决策与行动的机器人中心框架。该框架接收多模态输入,先感知环境与人类状态,再通过心智推理建模自我与他人的心理状态,最终基于推断出的心理状态生成决策与动作。此外,我们引入新型优化目标Mind-Reward,促使视觉语言模型产生一致的心智推理与行为表现。实验表明,该模型在决策与动作生成上分别优于GPT-4o 12.77%与12.49%。
原文摘要 · Abstract (English)
Theory of Mind (ToM) refers to the ability to infer others' mental states, such as beliefs, desires, and intentions. Current vision-language embodied agents lack ToM-based decision-making, and existing benchmarks focus solely on human mental states while ignoring the agent's own perspective, hindering coherent decision and action generation. To address this, we propose MindPower, a Robot-Centric framework integrating Perception, Mental Reasoning, Decision Making and Action. Given multimodal inputs, MindPower first perceives the environment and human states, then performs ToM Reasoning to model both self and others, and finally generates decisions and actions guided by inferred mental states. Furthermore, we introduce Mind-Reward, a novel optimization objective that encourages VLMs to produce consistent ToM Reasoning and behavior. Our model outperforms GPT-4o by 12.77% in decision making and 12.49% in action generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。