arXiv:2412.14803cs.CVcs.RO2024-12ICML被引 302

用视频预测模型提升机器人泛化能力,显著改善复杂操作成功率。

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

论文配图:Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
图 1 · 摘自论文原文
  • 基于视频扩散模型预测未来帧,构建含动态信息的视觉表征。
  • 在真实机器人任务中实现31.6%成功率提升,通用性基准上相对领先18.6%。
  • 适用于需要复杂操作与环境适应的机器人系统研发者。

视觉表征在构建通用机器人策略中起关键作用。以往的视觉编码器通常通过单图重建或双图对比学习预训练,倾向于捕捉静态信息,常忽略对具身任务至关重要的动态特性。近期,视频扩散模型(VDMs)展现出预测未来帧的能力,并体现出对物理世界的深刻理解。我们假设VDMs天然生成的视觉表征同时包含当前静态信息和预测的未来动态,可为机器人动作学习提供有效指导。基于此,我们提出视频预测策略(VPP),其通过条件于VDM内部预测的未来表征,学习隐式逆动力学模型。为更精准预测未来,我们在机器人数据集及互联网人类操作数据上微调预训练视频基础模型。实验表明,相比先前最先进方法,VPP在Calvin ABC-D泛化基准上取得18.6%的相对提升,在复杂真实世界灵巧操作任务中成功率提高31.6%。项目页面:https://video-prediction-policy.github.io

原文摘要 · Abstract (English)

Visual representations play a crucial role in developing generalist robotic policies. Previous vision encoders, typically pre-trained with single-image reconstruction or two-image contrastive learning, tend to capture static information, often neglecting the dynamic aspects vital for embodied tasks. Recently, video diffusion models (VDMs) demonstrate the ability to predict future frames and showcase a strong understanding of physical world. We hypothesize that VDMs inherently produce visual representations that encompass both current static information and predicted future dynamics, thereby providing valuable guidance for robot action learning. Based on this hypothesis, we propose the Video Prediction Policy (VPP), which learns implicit inverse dynamics model conditioned on predicted future representations inside VDMs. To predict more precise future, we fine-tune pre-trained video foundation model on robot datasets along with internet human manipulation data. In experiments, VPP achieves a 18.6\% relative improvement on the Calvin ABC-D generalization benchmark compared to the previous state-of-the-art, and demonstrates a 31.6\% increase in success rates for complex real-world dexterous manipulation tasks. Project page at https://video-prediction-policy.github.io

机器人视频生成泛化扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。