arXiv:2411.09153cs.CVcs.RO2024-11NeurIPS被引 61

用视频扩散模型学习机器人动作的隐式动态,提升操作精度。

VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation

  • 分两阶段训练:先预训练生成视觉轨迹,再用注意力适配器转为逆动力学模型。
  • 在CALVIN上比GR-1提升11.7%,在小数据集OXE上精度超9%。
  • 适合研究具身智能、机器人控制与世界模型融合的学者参考。

利用大规模视频数据训练视频生成模型,展现出理解复杂物理动态的巨大潜力,提示可借助多样化的机器人轨迹数据构建统一的、具备动态感知能力的模型以增强机器人操作能力。然而,由于可用机器人数据量相对较小,直接拟合数据而不考虑视觉观测与动作之间的关系,可能导致数据利用效率低下。为此,我们提出VidMan(用于机器人操作的视频扩散模型),采用受神经科学双过程理论启发的两阶段训练机制,以提升稳定性并优化数据利用效率。第一阶段,VidMan在Open X-Embodiment数据集(OXE)上预训练,通过视频去噪扩散方式预测未来视觉轨迹,使模型获得对环境动态的长期横向感知能力;第二阶段,引入灵活高效的逐层自注意力适配器,通过参数共享将VidMan转化为高效逆动力学模型,预测由隐式动态知识调制的动作。VidMan在CALVIN基准测试中超越最先进基线模型GR-1,实现11.7%的相对提升,并在小规模的OXE数据集上实现超过9%的精度提升。这些结果有力证明了世界模型能显著提高机器人动作预测的精度。代码与模型将公开。

原文摘要 · Abstract (English)

Recent advancements utilizing large-scale video data for learning video generation models demonstrate significant potential in understanding complex physical dynamics. It suggests the feasibility of leveraging diverse robot trajectory data to develop a unified, dynamics-aware model to enhance robot manipulation. However, given the relatively small amount of available robot data, directly fitting data without considering the relationship between visual observations and actions could lead to suboptimal data utilization. To this end, we propose VidMan (Video Diffusion for Robot Manipulation), a novel framework that employs a two-stage training mechanism inspired by dual-process theory from neuroscience to enhance stability and improve data utilization efficiency. Specifically, in the first stage, VidMan is pre-trained on the Open X-Embodiment dataset (OXE) for predicting future visual trajectories in a video denoising diffusion manner, enabling the model to develop a long horizontal awareness of the environment's dynamics. In the second stage, a flexible yet effective layer-wise self-attention adapter is introduced to transform VidMan into an efficient inverse dynamics model that predicts action modulated by the implicit dynamics knowledge via parameter sharing. Our VidMan framework outperforms state-of-the-art baseline model GR-1 on the CALVIN benchmark, achieving a 11.7% relative improvement, and demonstrates over 9% precision gains on the OXE small-scale dataset. These results provide compelling evidence that world models can significantly enhance the precision of robot action prediction. Codes and models will be public.

机器人操作扩散模型世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。