arXiv:2502.01784cs.ROcs.CV2025-02被引 10

用潜在视频规划提升机器人模仿学习的效率与一致性

VILP: Imitation Learning with Latent Video Planning

  • 基于潜在视频扩散模型生成多视角时序一致的机器人动作视频
  • 双视角各6帧、96x160分辨率视频生成达5赫兹,速度高效
  • 减少对高质量特定任务数据依赖,适合复杂多模态动作建模

在生成式AI时代,将视频生成模型融入机器人系统为通用机器人智能开辟了新路径。本文提出潜在视频规划模仿学习(VILP),构建一种潜在视频扩散模型,可生成具有良好时序一致性的预测机器人视频。该方法能从多视角生成高度时间对齐的视频,对机器人策略学习至关重要。其视频生成过程极高效:可在5赫兹速率下生成两个视角、每视角6帧、分辨率96x160像素的视频。实验表明,VILP在训练成本、推理速度、生成视频时序一致性及策略性能上均优于现有方法。相比其他模仿学习方法,VILP能以较少的高质量特定任务动作数据实现稳健表现,且具备强多模态动作分布表达能力。本工作为有效融合视频生成模型至机器人策略提供了实用范例,或对相关领域具启发意义。更多信息请见开源仓库:https://github.com/ZhengtongXu/VILP。

原文摘要 · Abstract (English)

In the era of generative AI, integrating video generation models into robotics opens new possibilities for the general-purpose robot agent. This paper introduces imitation learning with latent video planning (VILP). We propose a latent video diffusion model to generate predictive robot videos that adhere to temporal consistency to a good degree. Our method is able to generate highly time-aligned videos from multiple views, which is crucial for robot policy learning. Our video generation model is highly time-efficient. For example, it can generate videos from two distinct perspectives, each consisting of six frames with a resolution of 96x160 pixels, at a rate of 5 Hz. In the experiments, we demonstrate that VILP outperforms the existing video generation robot policy across several metrics: training costs, inference speed, temporal consistency of generated videos, and the performance of the policy. We also compared our method with other imitation learning methods. Our findings indicate that VILP can rely less on extensive high-quality task-specific robot action data while still maintaining robust performance. In addition, VILP possesses robust capabilities in representing multi-modal action distributions. Our paper provides a practical example of how to effectively integrate video generation models into robot policies, potentially offering insights for related fields and directions. For more details, please refer to our open-source repository https://github.com/ZhengtongXu/VILP.

模仿学习视频生成机器人扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。