用合成视频让机器人零样本学会模仿人类动作
EmbodiSwap for Zero-Shot Robot Imitation Learning
- 将人类视频与机器人模型结合,生成逼真合成数据
- 零样本训练模型在真实世界达82%成功率
- 适合做机器人模仿学习的快速原型设计
我们提出EmbodiSwap,一种在人类视频上生成逼真机器人合成影像的方法。该方法用于零样本机器人模仿学习,弥合了真实场景中第一人称人类视频与目标机器人形态之间的具身差异。我们基于EmbodiSwap生成的数据训练闭环机器人操作策略。首次将V-JEPA作为视觉主干网络,从视频理解领域迁移到合成机器人视频的模仿学习任务中。相比传统机器人常用视觉主干,V-JEPA表现更优。在真实测试中,零样本训练的V-JEPA模型达到82%成功率,优于少量样本训练的π₀网络以及在EmbodiSwap数据上训练的π₀网络。我们公开了:(i) 输入人类视频与任意机器人URDF即可生成机器人数据集的代码;(ii) 在EPIC-Kitchens、HOI4D和Ego4D上合成的机器人数据集;(iii) 模型检查点与推理代码,以支持可复现研究与广泛应用。
原文摘要 · Abstract (English)
We introduce EmbodiSwap - a method for producing photorealistic synthetic robot overlays over human video. We employ EmbodiSwap for zero-shot imitation learning, bridging the embodiment gap between in-the-wild ego-centric human video and a target robot embodiment. We train a closed-loop robot manipulation policy over the data produced by EmbodiSwap. We make novel use of V-JEPA as a visual backbone, repurposing V-JEPA from the domain of video understanding to imitation learning over synthetic robot videos. Adoption of V-JEPA outperforms alternative vision backbones more conventionally used within robotics. In real-world tests, our zero-shot trained V-JEPA model achieves an $82\%$ success rate, outperforming a few-shot trained $π_0$ network as well as $π_0$ trained over data produced by EmbodiSwap. We release (i) code for generating the synthetic robot overlays which takes as input human videos and an arbitrary robot URDF and generates a robot dataset, (ii) the robot dataset we synthesize over EPIC-Kitchens, HOI4D and Ego4D, and (iii) model checkpoints and inference code, to facilitate reproducible research and broader adoption.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。