arXiv:2505.20290cs.ROcs.AI2025-05被引 41

用智能眼镜拍摄的人类示范数据,零机器人数据训练出可通用的抓取策略。

EgoZero: Robot Learning from Smart Glasses

  • 仅用人类第一视角视频和零机器人数据训练
  • 7个任务零样本迁移成功率70%,每任务仅需20分钟数据
  • 适合想低成本获取真实世界机器人技能的研究者

尽管通用机器人技术取得进展,但实际应用中机器人策略仍远落后于人类基本能力。人类与物理世界的持续交互产生了丰富数据,却未被充分用于机器人学习。我们提出EgoZero,一个极简系统,仅依靠Project Aria智能眼镜采集的自然场景人类示范视频,以及零机器人数据,即可学习鲁棒的操控策略。EgoZero实现:(1) 从野外、第一人称视角的人类示范中提取完整可执行动作;(2) 将人类视觉观察压缩为与形态无关的状态表示;(3) 实现形态、空间和语义上的闭环策略泛化。我们在Franka Panda机械臂上部署策略,在7个操控任务上实现70%的零样本迁移成功率,每任务仅需20分钟数据收集。结果表明,野外人类数据可作为现实世界机器人学习的可扩展基础,推动机器人获得丰富、多样且自然的训练数据。代码与视频见https://egozero-robot.github.io。

原文摘要 · Abstract (English)

Despite recent progress in general purpose robotics, robot policies still lag far behind basic human capabilities in the real world. Humans interact constantly with the physical world, yet this rich data resource remains largely untapped in robot learning. We propose EgoZero, a minimal system that learns robust manipulation policies from human demonstrations captured with Project Aria smart glasses, $\textbf{and zero robot data}$. EgoZero enables: (1) extraction of complete, robot-executable actions from in-the-wild, egocentric, human demonstrations, (2) compression of human visual observations into morphology-agnostic state representations, and (3) closed-loop policy learning that generalizes morphologically, spatially, and semantically. We deploy EgoZero policies on a gripper Franka Panda robot and demonstrate zero-shot transfer with 70% success rate over 7 manipulation tasks and only 20 minutes of data collection per task. Our results suggest that in-the-wild human data can serve as a scalable foundation for real-world robot learning - paving the way toward a future of abundant, diverse, and naturalistic training data for robots. Code and videos are available at https://egozero-robot.github.io.

机器人学习第一视角零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。