用人类第一视角视频训练人形机器人,实现复杂环境下的自主操作。
EgoHumanoid: Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration
- 通过视觉-语言-动作联合训练,融合大量人类第一视角数据与少量机器人数据。
- 在未知环境中性能比纯机器人数据训练提升51%,显著增强泛化能力。
- 提出视角与动作对齐机制,解决人体与机器人形态差异带来的迁移难题。
人类示范具备丰富的环境多样性和自然扩展性,是机器人遥操作的有力替代方案。尽管该范式已在机械臂操作中取得进展,其在更具挑战性、数据需求量大的人形机器人运动-操作协同任务中的潜力仍待探索。本文提出EgoHumanoid,首个利用海量第一视角人类示范与少量机器人数据共同训练视觉-语言-动作策略的框架,使机器人能在多样真实环境中完成运动-操作任务。为弥合人类与机器人之间的具身差距(包括物理形态和视角差异),我们构建了从硬件设计到数据处理的系统性对齐流程。开发了可扩展的人类数据采集便携系统,并建立实用采集规范以提升迁移效果。核心对齐管道包含两项关键组件:视角对齐减少因摄像机高度与视角变化引起的视觉域差异;动作对齐将人类动作映射至统一且运动学可行的人形控制空间。大量真实世界实验表明,引入无机器人参与的第一视角数据,使性能相比仅使用机器人数据的基线提升51%,尤其在未见环境中表现更优。分析进一步揭示了可有效迁移的行为模式及人类数据规模化的潜力。
原文摘要 · Abstract (English)
Human demonstrations offer rich environmental diversity and scale naturally, making them an appealing alternative to robot teleoperation. While this paradigm has advanced robot-arm manipulation, its potential for the more challenging, data-hungry problem of humanoid loco-manipulation remains largely unexplored. We present EgoHumanoid, the first framework to co-train a vision-language-action policy using abundant egocentric human demonstrations together with a limited amount of robot data, enabling humanoids to perform loco-manipulation across diverse real-world environments. To bridge the embodiment gap between humans and robots, including discrepancies in physical morphology and viewpoint, we introduce a systematic alignment pipeline spanning from hardware design to data processing. A portable system for scalable human data collection is developed, and we establish practical collection protocols to improve transferability. At the core of our human-to-humanoid alignment pipeline lies two key components. The view alignment reduces visual domain discrepancies caused by camera height and perspective variation. The action alignment maps human motions into a unified, kinematically feasible action space for humanoid control. Extensive real-world experiments demonstrate that incorporating robot-free egocentric data significantly outperforms robot-only baselines by 51\%, particularly in unseen environments. Our analysis further reveals which behaviors transfer effectively and the potential for scaling human data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。