从网络视频中蒸馏机器人操作技能,无需真实机器人演示。
ZeroMimic: Distilling Robotic Manipulation Skills from Web Videos
- 利用人类视角视频中的语义与几何信息,自动提取操作技能。
- 在未见过的厨房场景中,直接部署6类操作任务策略,表现稳定。
- 适合希望快速复用通用操作能力的研究者与开发者。
近期机器人操作进展多依赖模仿学习,但通常需在相同机器人、环境和物体下收集演示数据,难以获取。相比之下,大量展示真实世界操作行为的人类视频数据已存在,蕴含丰富可用信息。能否仅凭这些视频,不依赖额外机器人演示或探索,蒸馏出可立即部署的机器人技能策略?我们提出首个实现此目标的系统 ZeroMimic,能生成针对常见操作任务(开、关、倒、拾取放置、切割、搅拌)的图像目标条件化技能策略,每类策略可适配多种物体及未见任务布局。ZeroMimic 融合最新视觉理解技术、现代抓取可操作性检测器与模仿策略架构,在 EpicKitchens 人类视角视频数据集上训练后,直接在真实与仿真厨房环境中验证其泛化能力,使用两种不同机器人形态进行测试。为支持即插即用复用,我们发布完整软件与策略检查点。
原文摘要 · Abstract (English)
Many recent advances in robotic manipulation have come through imitation learning, yet these rely largely on mimicking a particularly hard-to-acquire form of demonstrations: those collected on the same robot in the same room with the same objects as the trained policy must handle at test time. In contrast, large pre-recorded human video datasets demonstrating manipulation skills in-the-wild already exist, which contain valuable information for robots. Is it possible to distill a repository of useful robotic skill policies out of such data without any additional requirements on robot-specific demonstrations or exploration? We present the first such system ZeroMimic, that generates immediately deployable image goal-conditioned skill policies for several common categories of manipulation tasks (opening, closing, pouring, pick&place, cutting, and stirring) each capable of acting upon diverse objects and across diverse unseen task setups. ZeroMimic is carefully designed to exploit recent advances in semantic and geometric visual understanding of human videos, together with modern grasp affordance detectors and imitation policy classes. After training ZeroMimic on the popular EpicKitchens dataset of ego-centric human videos, we evaluate its out-of-the-box performance in varied real-world and simulated kitchen settings with two different robot embodiments, demonstrating its impressive abilities to handle these varied tasks. To enable plug-and-play reuse of ZeroMimic policies on other task setups and robots, we release software and policy checkpoints of our skill policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。