从真实视频中学习可迁移的动态知识,提升机器人长时任务执行能力。
VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
- 用动态增强的潜在动力学模型分离动作与视觉,专注学习任务相关动态。
- 在真实手工制作任务中任务成功率提升70%,生成连贯长视频。
- 适合研究视频理解、机器人操作与世界模型的学者使用。
从无标签视频数据中学习可迁移知识并应用于新环境,是智能体的核心能力。本文提出VideoWorld 2,首次直接从原始真实视频中学习可迁移知识。其核心是动态增强的潜在动力学模型(dLDM),将动作动态与视觉外观解耦:预训练的视频扩散模型负责视觉建模,使dLDM专注于学习紧凑且有意义的任务相关潜在代码。这些潜在代码被自回归建模以学习任务策略,并支持长时推理。我们在具有挑战性的真实世界手工制作任务上评估VideoWorld 2,此前的视频生成和潜在动力学模型在此类任务中表现不稳定。结果表明,VideoWorld 2任务成功率最高提升70%,并能生成连贯的长时执行视频。在机器人领域,我们证明其可从Open-X数据集获取有效操作知识,显著提升CALVIN上的任务表现。本研究揭示了直接从原始视频中学习可迁移世界知识的潜力,所有代码、数据与模型将开源,供进一步研究。
原文摘要 · Abstract (English)
Learning transferable knowledge from unlabeled video data and applying it in new environments is a fundamental capability of intelligent agents. This work presents VideoWorld 2, which extends VideoWorld and offers the first investigation into learning transferable knowledge directly from raw real-world videos. At its core, VideoWorld 2 introduces a dynamic-enhanced Latent Dynamics Model (dLDM) that decouples action dynamics from visual appearance: a pretrained video diffusion model handles visual appearance modeling, enabling the dLDM to learn latent codes that focus on compact and meaningful task-related dynamics. These latent codes are then modeled autoregressively to learn task policies and support long-horizon reasoning. We evaluate VideoWorld 2 on challenging real-world handcraft making tasks, where prior video generation and latent-dynamics models struggle to operate reliably. Remarkably, VideoWorld 2 achieves up to 70% improvement in task success rate and produces coherent long execution videos. In robotics, we show that VideoWorld 2 can acquire effective manipulation knowledge from the Open-X dataset, which substantially improves task performance on CALVIN. This study reveals the potential of learning transferable world knowledge directly from raw videos, with all code, data, and models to be open-sourced for further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。