arXiv:2506.22355cs.AI2025-06被引 83

让智能体像人一样感知世界,通过建模实现自主决策与协作。

Embodied AI Agents: Modeling the World

  • 构建物理与心理双重世界模型,支撑智能体感知与推理。
  • 使虚拟化身、可穿戴设备和机器人能自主理解环境与用户意图。
  • 适合研究具身智能、人机协同的学者与工程师参考。

本文探讨了嵌入视觉、虚拟或物理形态中的人工智能智能体,使其能够与用户及环境互动。这些智能体包括虚拟化身、可穿戴设备和机器人,设计目标是实现对周围环境的感知、学习与行动,从而更接近人类的学习与交互方式,相较于无体智能体更具优势。我们提出,世界建模是具身智能体推理与规划的核心,有助于理解并预测环境,识别用户意图与社会情境,进而提升其自主执行复杂任务的能力。世界建模融合多模态感知、基于推理的动作与控制规划,以及记忆机制,以形成对物理世界的全面认知。此外,我们还提出学习用户的“心智世界模型”,以促进更高效的人机协作。

原文摘要 · Abstract (English)

This paper describes our research on AI agents embodied in visual, virtual or physical forms, enabling them to interact with both users and their environments. These agents, which include virtual avatars, wearable devices, and robots, are designed to perceive, learn and act within their surroundings, which makes them more similar to how humans learn and interact with the environments as compared to disembodied agents. We propose that the development of world models is central to reasoning and planning of embodied AI agents, allowing these agents to understand and predict their environment, to understand user intentions and social contexts, thereby enhancing their ability to perform complex tasks autonomously. World modeling encompasses the integration of multimodal perception, planning through reasoning for action and control, and memory to create a comprehensive understanding of the physical world. Beyond the physical world, we also propose to learn the mental world model of users to enable better human-agent collaboration.

具身智能世界模型人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。