用人类第一视角视频零样本训练机器人,30分钟搞定真实任务
HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos

- 将人手-物体交互抽象为实体级表示,增强动作模仿信号
- 仅需30分钟人类视频,4个任务平均成功率92.5%,15分钟达75%
- 无需机器人数据,可零样本迁移至新机器人、摄像头和环境
人类第一视角视频无需机器人硬件即可捕获丰富操作示范,但因人体与机器人在视觉外观和运动学上的具身差异,技能迁移仍具挑战。我们提出HumanEgo框架,通过将每段人类示范提升为手-物交互的实体级表示,并使用密集辅助目标训练流匹配策略,强化每条轨迹的监督信号。HumanEgo不依赖机器人数据,硬件无关,数据高效,支持零样本的人类到机器人的技能迁移。每项任务仅需30分钟人类视频,即可在四个真实任务中实现92.5%的平均成功率(15分钟时达75%),比等时机器人遥操作高41%,且能鲁棒地零样本迁移到新机器人、摄像头和环境中。我们已开源该框架,便于直接从人类数据学习机器人策略:https://github.com/TX-Leo/HumanEgo
原文摘要 · Abstract (English)
Human egocentric video captures rich manipulation demonstrations without any robot hardware, yet transferring these skills to robots remains challenging due to the embodiment gap between human and robot in both visual appearance and kinematics. We present HumanEgo, a framework that bridges the embodiment gap by lifting each human demonstration to an entity-level representation of hand-object interaction, and training a flow matching policy with dense auxiliary objectives that amplify supervision from every trajectory. HumanEgo is robot-data-free, hardware-agnostic, data-efficient, and zero-shot human-to-robot transferable. With only 30 minutes of human videos per task, HumanEgo achieves 92.5% average success across four real-world tasks (75% with just 15 minutes), outperforms matched-time robot teleoperation by 41%, and robustly transfers zero-shot across novel robots, cameras, and environments. We release HumanEgo as an easy-to-use, open-source framework for learning robot policies directly from human data: https://github.com/TX-Leo/HumanEgo
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。