用真实机器人数据预训练动态模型,让机器人学会通用交互规律。
Learning Transferable Dynamics Priors from Action to World Modeling

- 用多视角扩散模型学习动作如何改变视觉场景
- 预训练模型可转为仿真器或策略预测器,支持长时序推演
- 适合需要快速模拟和策略学习的机器人研究者
我们研究了以动作条件的世界建模,作为学习可迁移动力学先验的可扩展方式。通过在带有真实动作标注的大规模机器人操作数据上预训练一个多视角交互式基础扩散世界模型A2World,该模型能够捕捉超越外观级视频生成的可重用交互动态。我们从两个互补角度验证了所学动力学先验:首先,将A2World适配为任务或场景特化的现实世界仿真器A2World-sim,其长时序推演可替代真实机器人推演,支持基于仿真的策略评估与大规模假设分析;其次,从相同预训练权重出发,将A2World适配为视频-动作联合预测模型A2World-policy,实现视觉与指令条件下的动作预测。在仿真基准和真实机器人设置中的实验表明,动作条件的世界模型预训练能产生对仿真驱动与策略驱动机器人学习均有帮助的可迁移动力学先验。
原文摘要 · Abstract (English)
We study action-conditioned world modeling as a scalable way to learn transferable dynamics priors for robot learning. By pretraining a model to predict how actions drive visual scene evolution, the resulting world model captures reusable interaction dynamics beyond appearance-level video generation. Concretely, we pretrain a multi-view interactive base diffusion world model, A2World, on large-scale robot manipulation data with real action annotations. We validate the learned dynamics priors from two complementary perspectives. First, we adapt A2World into a task- or scene-specialized real-world simulator, A2World-sim, whose long-horizon rollouts support simulator-based policy evaluation and scalable what-if analysis by replacing real-robot rollouts with world model rollouts. Second, starting from the same pretrained weights, we adapt A2World into a video-action joint prediction model, A2World-policy, that predicts actions under visual and instruction conditioning. Experiments across simulation benchmarks and real-robot settings demonstrate that action-conditioned world model pretraining yields transferable dynamics priors that benefit both simulator-centric and policy-centric robot learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。