用视频预测未来,让机器人零样本学会新动作。
World Action Models are Zero-shot Policies
- 基于视频扩散模型预测世界状态与动作,学习物理动态。
- 真实机器人实验中任务泛化能力提升2倍以上。
- 仅需10-20分钟视频数据即可跨机器人迁移,30分钟适配新身体。
当前最先进的视觉-语言-动作(VLA)模型在语义泛化上表现优异,但在新环境中的未见物理动作泛化能力不足。我们提出DreamZero,一种基于预训练视频扩散主干的的世界动作模型(WAM)。与VLA不同,WAM通过预测未来世界状态和动作来学习物理动态,将视频作为世界演化的密集表征。通过联合建模视频与动作,DreamZero能从异构机器人数据中有效学习多样化技能,无需重复演示。在真实机器人实验中,其对新任务和环境的泛化能力相比最优VLA提升超过2倍。关键的是,通过模型与系统优化,我们使140亿参数的自回归视频扩散模型实现7Hz的实时闭环控制。最后,我们展示了两种跨体感迁移:其他机器人或人类的纯视频示范,仅需10-20分钟数据即可带来超42%的任务性能提升;更令人意外的是,仅用30分钟游戏数据,DreamZero即可完成少样本体感适应,同时保持零样本泛化能力。
原文摘要 · Abstract (English)
State-of-the-art Vision-Language-Action (VLA) models excel at semantic generalization but struggle to generalize to unseen physical motions in novel environments. We introduce DreamZero, a World Action Model (WAM) built upon a pretrained video diffusion backbone. Unlike VLAs, WAMs learn physical dynamics by predicting future world states and actions, using video as a dense representation of how the world evolves. By jointly modeling video and action, DreamZero learns diverse skills effectively from heterogeneous robot data without relying on repetitive demonstrations. This results in over 2x improvement in generalization to new tasks and environments compared to state-of-the-art VLAs in real robot experiments. Crucially, through model and system optimizations, we enable a 14B autoregressive video diffusion model to perform real-time closed-loop control at 7Hz. Finally, we demonstrate two forms of cross-embodiment transfer: video-only demonstrations from other robots or humans yield a relative improvement of over 42% on unseen task performance with just 10-20 minutes of data. More surprisingly, DreamZero enables few-shot embodiment adaptation, transferring to a new embodiment with only 30 minutes of play data while retaining zero-shot generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。