仅用一段视频教人形机器人完成操作,还能适应不同物体位置。
OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation

- 通过感知物体并分别调整身体与手部动作,实现视频中人类动作的精准模仿。
- 在多种视觉和空间条件下表现优秀,成功率达79.2%,无需人工遥控。
- 适合希望快速部署机器人操作技能的研究者与工程师。
我们研究如何通过单段视频示范,教人人形机器人掌握操作技能。提出OKAMI方法,从单个RGB-D视频生成操作规划并推导执行策略。核心是物体感知的重定向技术,使机器人能在部署时根据物体位置变化,模仿视频中的人类动作。OKAMI利用开放世界视觉模型识别任务相关物体,并分别重定向身体运动与手部姿态。实验表明,OKAMI在多种视觉与空间条件下表现出强泛化能力,优于当前最先进的基于观察的开放世界模仿学习方法。此外,利用OKAMI生成的轨迹训练闭环视觉-运动策略,平均成功率高达79.2%,无需繁琐的人工遥操作。更多演示视频见官网:https://ut-austin-rpl.github.io/OKAMI/
原文摘要 · Abstract (English)
We study the problem of teaching humanoid robots manipulation skills by imitating from single video demonstrations. We introduce OKAMI, a method that generates a manipulation plan from a single RGB-D video and derives a policy for execution. At the heart of our approach is object-aware retargeting, which enables the humanoid robot to mimic the human motions in an RGB-D video while adjusting to different object locations during deployment. OKAMI uses open-world vision models to identify task-relevant objects and retarget the body motions and hand poses separately. Our experiments show that OKAMI achieves strong generalizations across varying visual and spatial conditions, outperforming the state-of-the-art baseline on open-world imitation from observation. Furthermore, OKAMI rollout trajectories are leveraged to train closed-loop visuomotor policies, which achieve an average success rate of 79.2% without the need for labor-intensive teleoperation. More videos can be found on our website https://ut-austin-rpl.github.io/OKAMI/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。