arXiv:2410.24221cs.ROcs.CV2024-10ICRA被引 226

用第一视角视频和手部追踪数据,让机器人学会像人一样操作物体。

EgoMimic: Scaling Imitation Learning via Egocentric Video

论文配图:EgoMimic: Scaling Imitation Learning via Egocentric Video
图 1 · 摘自论文原文
  • 用智能眼镜采集人类第一视角操作视频与3D手部动作数据
  • 在长时序任务上超越现有方法,支持新场景泛化
  • 增加1小时人手数据比1小时机器人数据更有效

模仿学习所需演示数据的规模与多样性是重大挑战。本文提出EgoMimic,一个全栈式框架,通过人类具身数据(特别是第一视角视频与3D手部追踪)实现操纵技能的规模化学习。该框架包含:(1) 使用人体工学Project Aria眼镜采集人类具身数据;(2) 低成本双臂机械臂以最小化与人类动作的运动学差距;(3) 跨域数据对齐技术;(4) 联合训练人类与机器人数据的模仿学习架构。相比仅提取人类视频高层意图的方法,EgoMimic将人类与机器人数据视为同等的具身示范数据,学习统一策略。实验表明,EgoMimic在多种长时序、单臂及双臂操作任务中显著优于现有先进方法,并实现对全新场景的泛化能力。最后,我们验证了其良好的可扩展性:增加1小时的人类手部数据带来的收益远高于1小时机器人数据。

原文摘要 · Abstract (English)

The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos paired with 3D hand tracking. EgoMimic achieves this through: (1) a system to capture human embodiment data using the ergonomic Project Aria glasses, (2) a low-cost bimanual manipulator that minimizes the kinematic gap to human data, (3) cross-domain data alignment techniques, and (4) an imitation learning architecture that co-trains on human and robot data. Compared to prior works that only extract high-level intent from human videos, our approach treats human and robot data equally as embodied demonstration data and learns a unified policy from both data sources. EgoMimic achieves significant improvement on a diverse set of long-horizon, single-arm and bimanual manipulation tasks over state-of-the-art imitation learning methods and enables generalization to entirely new scenes. Finally, we show a favorable scaling trend for EgoMimic, where adding 1 hour of additional hand data is significantly more valuable than 1 hour of additional robot data. Videos and additional information can be found at https://egomimic.github.io/

模仿学习第一视角具身智能机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。