arXiv:2607.11689cs.ROcs.AI2026-07被引 1

构建可自适应的具身智能体,让机器像人一样在真实世界中推理与行动。

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

论文配图:From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence
图 1 · 摘自论文原文
  • 以具身大脑为核心,整合多模态信息并评估干预方案
  • 提出闭环训练机制,将验证后的交互经验转化为可复用知识
  • 通过统一接口连接异构模型与任务,推动物理智能系统模块化

通用人工智能最终需要能够在物理世界中推理与行动的智能体。动作模型、视觉-语言-动作策略和世界模型虽已取得进展,但世界动作模型(WAMs)尤为关键,因其能将干预行为与预测后果相联系。然而当前进展仍零散:模型使用不兼容的动作空间与预测目标,数据集与任务标准各异,运行系统接口有限,难以复用与评估。本文梳理了向WAMs演进的过程,并归纳出三大耦合差距:模型角色与表征、目标与标准化、系统组合。基于此,提出一个协同发展路线图,聚焦于‘具身大脑’——一种长期目标模型,能融合多模态上下文,比较候选干预方案,并发出状态转移或能力请求,而非直接控制执行器命令。WAMs为其实现预测功能提供了可行原型,而物理约束装置通过工具、控制器、验证与日志记录使模型输出落地。共享合约协调异构模型、数据、任务与具身形式,闭环后训练将已验证的交互转化为可复用经验。这些组件共同定义了一个模块化的物理智能栈,支持自适应与自我提升的具身智能体。

原文摘要 · Abstract (English)

Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) are particularly promising because they connect candidate interventions with predicted consequences. However, progress remains fragmented: models use incompatible action spaces and prediction targets, datasets and tasks follow different conventions, and runtime systems expose limited interfaces for reuse and evaluation. We review the evolution toward WAMs and organize these limitations into three coupled gaps: model roles and representations, objectives and standardization, and system composition. Building on this analysis, we propose a co-evolution roadmap for physical intelligence centered on the \emph{embodied brain}, a long-term model target for integrating multimodal context, comparing candidate interventions, and issuing state-transition or capability requests rather than direct actuator commands. WAMs provide promising prototypes for its predictive functions, while a physical harness grounds model outputs through tools, controllers, verification, and trace logging. Shared contracts align heterogeneous models, data, tasks, and embodiments, and closed-loop post-training converts verified interaction into reusable experience. Together, these components define a modular physical-intelligence stack for adaptive and self-improving embodied agents.

具身智能世界模型自适应系统物理智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。