仅用人类视频训练机器人,无需真实机器人数据即可直接部署。
Phantom: Training Robots Without Robots Using Only Human Videos
- 用人体姿态估计和视觉编辑将人类视频转为机器人可用的观测-动作对。
- 在多种任务上实现最高92%的成功率,包括柔性物体操作和多物清扫。
- 零样本部署到真实硬件,适合希望低成本训练通用机器人的研究者。
训练通用机器人需要从大量多样数据中学习。当前方法严重依赖遥操作示范,难以扩展。本文提出一种可扩展框架,直接从人类视频示范训练抓取策略,无需任何机器人数据。通过人体姿态估计和视觉编辑,将人类示范转化为机器人兼容的观测-动作对,擦除人臂并叠加渲染机器人以对齐视觉域。该方法支持零样本部署至真实硬件,无需微调。我们在多种任务上验证了其有效性,成功率高达92%,涵盖柔性物体操作、多物体清扫及插入任务。方法具备对新环境的泛化能力,并支持闭环执行。本工作证明仅凭人类视频即可训练有效策略,为可扩展机器人学习开辟新路径。
原文摘要 · Abstract (English)
Training general-purpose robots requires learning from large and diverse data sources. Current approaches rely heavily on teleoperated demonstrations which are difficult to scale. We present a scalable framework for training manipulation policies directly from human video demonstrations, requiring no robot data. Our method converts human demonstrations into robot-compatible observation-action pairs using hand pose estimation and visual data editing. We inpaint the human arm and overlay a rendered robot to align the visual domains. This enables zero-shot deployment on real hardware without any fine-tuning. We demonstrate strong success rates-up to 92%-on a range of tasks including deformable object manipulation, multi-object sweeping, and insertion. Our approach generalizes to novel environments and supports closed-loop execution. By demonstrating that effective policies can be trained using only human videos, our method broadens the path to scalable robot learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。