arXiv:2507.00990cs.ROcs.AI2025-07被引 44

无需真人示范,机器人靠模仿AI生成视频完成复杂操作。

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations

  • 用语言和初始画面生成演示视频,再筛选符合指令的视频
  • 生成视频经筛选后效果接近真实示范,且质量越高越有效
  • 适合想快速部署新任务的机器人研发者,无需训练数据

本文提出机器人模仿生成视频(RIGVid)系统,使机器人仅通过模仿AI生成的视频即可完成倒水、擦拭、搅拌等复杂操作,无需任何物理示范或机器人专属训练。给定语言指令和初始场景图像,视频扩散模型生成潜在示范视频,视觉语言模型(VLM)自动过滤不符合指令的结果。6D姿态追踪器从视频中提取物体轨迹,并以与机体无关的方式重定向至机器人。大量真实世界测试表明,经筛选的生成视频效果与真实示范相当,且性能随生成质量提升而增强。相比基于VLM的关键点预测等紧凑替代方案,生成视频表现更优;强6D姿态追踪也优于密集特征点跟踪。结果表明,由先进现成模型生成的视频可作为机器人操作的有效监督来源。

原文摘要 · Abstract (English)

This work introduces Robots Imitating Generated Videos (RIGVid), a system that enables robots to perform complex manipulation tasks--such as pouring, wiping, and mixing--purely by imitating AI-generated videos, without requiring any physical demonstrations or robot-specific training. Given a language command and an initial scene image, a video diffusion model generates potential demonstration videos, and a vision-language model (VLM) automatically filters out results that do not follow the command. A 6D pose tracker then extracts object trajectories from the video, and the trajectories are retargeted to the robot in an embodiment-agnostic fashion. Through extensive real-world evaluations, we show that filtered generated videos are as effective as real demonstrations, and that performance improves with generation quality. We also show that relying on generated videos outperforms more compact alternatives such as keypoint prediction using VLMs, and that strong 6D pose tracking outperforms other ways to extract trajectories, such as dense feature point tracking. These findings suggest that videos produced by a state-of-the-art off-the-shelf model can offer an effective source of supervision for robotic manipulation.

机器人操作视频生成零样本学习模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。