用4.4万小时人类视频训练通用机器人世界模型,实现高精度动作控制。
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

- 从4.4万小时第一视角视频中学习多样化交互与精细操控。
- 在小规模机器人数据微调后,物理理解与动作可控性显著提升。
- 支持实时推理(10.81 FPS)和远程操作等实际应用。
能够模拟不同环境下动作结果的能力将推动通用智能体的大规模发展。然而,由于数据覆盖有限且动作标签稀缺,建模这些世界动态,特别是灵巧机器人任务,仍面临重大挑战。为此,我们提出DreamDojo,一个基于44,000小时第一人称人类视频的通用世界模型基础模型。该数据混合是迄今最大规模的世界模型预训练视频数据集,涵盖广泛日常场景、多样物体与技能。为应对动作标签稀缺问题,我们引入连续潜在动作作为统一代理动作,增强未标注视频中的交互知识迁移。在小规模目标机器人数据上微调后,DreamDojo展现出对物理规律的深刻理解与精确的动作可控性。我们还设计了一种蒸馏流水线,使模型达到10.81 FPS的实时推理速度,并进一步提升上下文一致性。本工作实现了基于生成式世界模型的重要应用,包括实时远程操作、策略评估与模型基于规划。在多个具有挑战性的分布外(OOD)基准上的系统评估验证了方法在开放世界、接触密集任务中的有效性,为通用机器人世界模型的发展铺平道路。
原文摘要 · Abstract (English)
Being able to simulate the outcomes of actions in varied environments will revolutionize the development of generalist agents at scale. However, modeling these world dynamics, especially for dexterous robotics tasks, poses significant challenges due to limited data coverage and scarce action labels. As an endeavor towards this end, we introduce DreamDojo, a foundation world model that learns diverse interactions and dexterous controls from 44k hours of egocentric human videos. Our data mixture represents the largest video dataset to date for world model pretraining, spanning a wide range of daily scenarios with diverse objects and skills. To address the scarcity of action labels, we introduce continuous latent actions as unified proxy actions, enhancing interaction knowledge transfer from unlabeled videos. After post-training on small-scale target robot data, DreamDojo demonstrates a strong understanding of physics and precise action controllability. We also devise a distillation pipeline that accelerates DreamDojo to a real-time speed of 10.81 FPS and further improves context consistency. Our work enables several important applications based on generative world models, including live teleoperation, policy evaluation, and model-based planning. Systematic evaluation on multiple challenging out-of-distribution (OOD) benchmarks verifies the significance of our method for simulating open-world, contact-rich tasks, paving the way for general-purpose robot world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。