厘清世界动作模型边界,梳理预测与行动融合的设计范式。
World Action Models: A Survey
- 按生成内容分为渲染未来、潜在未来与无视频推理三类
- 揭示模型设计在计算开销与控制精度间的权衡规律
- 适合研究具身智能与交互式预测的开发者参考
世界动作模型(WAMs)是将未来预测结果直接用于行动决策的具身预测-动作模型。近期工作或复用大型视频生成模型,或基于语言/视觉-语言主干网络而无需视频生成核心,导致世界模型、视频生成模型、动作引导视频世界模型、视觉-语言-动作策略与WAMs之间的界限模糊。本综述为该领域建立统一认知:首先明确各类模型的边界,再从两个互补视角组织现有方法。第一视角关注模型需生成的内容,涵盖渲染未来、潜在未来和免视频生成的动作推理;第二视角分解为预测底座、主干网络、动作耦合与部署模式。该分析框架支持对可交互性、因果性、持续性、物理合理性与泛化能力的统一讨论,并进一步涵盖数据、评估与开放挑战。多维度分析揭示一致设计模式:WAMs并非仅带动作头的视频生成器,而是以表征丰富度换取算力、内存、延迟与动作标签成本的预测-动作系统。领域正朝着生成更少未来但保留必要控制信息的方向演进。相关主页见 https://world-action-models.github.io/。
原文摘要 · Abstract (English)
World Action Models (WAMs) are embodied predictive-action models that make a forecast of the future available to action. Recent WAMs repurpose large video generation models, and a parallel line relies on language or vision-language backbones without a video-generation core. This rapid expansion has blurred the boundary among broad world models, video generation models, action-grounded video world models, Vision-Language-Action policies, and WAMs. This survey gives the field a common account. It first clarifies these boundaries, then organizes existing works through two complementary views. The first view asks what each method is required to generate, spanning rendered futures, latent futures, and video-generation-free action reasoning. The second view decomposes each method by predictive substrate, backbone, action coupling, and deployment regime. This anatomy supports a unified discussion of interactability, causality, persistence, physical plausibility, and generalization, followed by data, evaluation, and open challenges. Across these axes, a consistent design pattern emerges: WAMs are not simply video generators with action heads, but predictive-action methods whose design choices trade representational richness against compute, memory, latency, and action-label cost. The field is moving toward methods that generate less of the future while preserving what control requires. The survey homepage is available at https://world-action-models.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。