arXiv:2605.12090cs.ROcs.CL2026-05被引 36

将世界模型与动作生成结合,让AI理解干预后世界如何变化。

World Action Models: The Next Frontier in Embodied AI

论文配图:World Action Models: The Next Frontier in Embodied AI
图 1 · 摘自论文原文
  • 构建统一框架,区分世界动作模型与传统方法
  • 归纳出级联与联合两类架构,涵盖生成模态与条件机制
  • 适合研究具身智能、机器人决策的学者参考

视觉-语言-动作(VLA)模型在具身策略学习中实现了强大的语义泛化,但仅学习观测到动作的反应映射,未显式建模干预下物理世界的演化。越来越多工作通过将世界模型——即环境动态的预测模型——引入动作生成流程来弥补这一缺陷。我们提出这一新兴范式为世界动作模型(WAMs):统一预测状态与动作生成的具身基础模型,旨在建模未来状态与动作的联合分布,而非单一动作。然而,现有研究在架构、学习目标和应用场景上仍碎片化,缺乏统一概念框架。本文正式定义了WAMs,厘清其与相关概念的区别,并追溯了VLA与世界模型研究融合的根源。我们将现有方法系统归类为级联与联合WAMs,进一步按生成模态、条件机制和动作解码策略细分。系统分析了推动WAMs发展的数据生态,包括机器人遥操作、便携人类示范、仿真环境及互联网规模的第一视角视频,并综述了围绕视觉保真度、物理常识与动作合理性展开的新兴评估协议。总体而言,本综述首次系统梳理了WAMs领域,阐明关键架构范式及其权衡,识别开放挑战与未来机遇。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have achieved strong semantic generalization for embodied policy learning, yet they learn reactive observation-to-action mappings without explicitly modeling how the physical world evolves under intervention. A growing body of work addresses this limitation by integrating world models, predictive models of environment dynamics, into the action generation pipeline. We term this emerging paradigm World Action Models (WAMs): embodied foundation models that unify predictive state modeling with action generation, targeting a joint distribution over future states and actions rather than actions alone. However, the literature remains fragmented across architectures, learning objectives, and application scenarios, lacking a unified conceptual framework. We formally define WAMs and disambiguate them from related concepts, and trace the foundations and early integration of VLA and world model research that gave rise to this paradigm. We organize existing methods into a structured taxonomy of Cascaded and Joint WAMs, with further subdivision by generation modality, conditioning mechanism, and action decoding strategy. We systematically analyze the data ecosystem fueling WAMs development, spanning robot teleoperation, portable human demonstrations, simulation, and internet-scale egocentric video, and synthesize emerging evaluation protocols organized around visual fidelity, physical commonsense, and action plausibility. Overall, this survey provides the first systematic account of the WAMs landscape, clarifies key architectural paradigms and their trade-offs, and identifies open challenges and future opportunities for this rapidly evolving field.

具身智能世界模型动作生成机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。