arXiv:2606.18772cs.RO2026-06

用真人示范训练人形机器人,实现精准动态抓取与全身协调。

HALOMI: Learning Humanoid Loco-Manipulation with Active Perception from Human Demonstrations

论文配图:HALOMI: Learning Humanoid Loco-Manipulation with Active Perception from Human Demonstrations
图 1 · 摘自论文原文
  • 从真人视角和手腕视角采集动作数据,构建自适应感知框架。
  • 在真实任务中平均成功率85%,支持抛掷与深蹲抓取等复杂动作。
  • 适合研究人形机器人操控、具身智能及示范学习的开发者。

真人示范可大规模采集,自然体现手眼协同,是学习人形机器人运动操作的潜在数据来源。然而,直接将真人示范迁移到人形机器人需依赖精确的世界坐标系追踪控制器,该控制器在分布外目标下往往不稳定,且真人与机器人在自我视角观测和动作执行上仍存在差距。为此,我们提出HALOMI,一个基于真人示范的可扩展人形运动操作学习框架。HALOMI在通用操作接口(UMI)基础上引入自我视角感知,规模化采集自我视角与手腕视角观测数据及头部-手部轨迹。我们进一步设计了一种流形约束控制器,在学习到的潜在行为流形中进行规划,实现世界坐标系下的精准鲁棒头手追踪。为弥合真人与机器人间的差异,我们实施自我视角对齐,并引入控制器感知的参考轨迹适配,降低观测与动作执行的偏差。我们在配备主动颈部的Unitree G1人形机器人上验证了HALOMI,涵盖导航、抓取、双臂操作、全身协调和动态行为等五项真实世界任务。在三项量化评估任务中,HALOMI平均成功率达85%;额外定性演示表明其具备动态抛掷与深蹲抓取能力。

原文摘要 · Abstract (English)

Human demonstrations, which can be collected at scale and naturally capture active hand-eye coordination, are a promising data source for learning humanoid loco-manipulation. However, directly transferring human demonstrations to humanoids requires a precise world-frame tracking controller, which is often brittle under Out-of-Distribution(OOD) targets, while human-to-humanoid gaps persist in both egocentric observation and action execution. To address these challenges, we present HALOMI, a scalable framework for learning humanoid loco-manipulation with active perception from human demonstrations. HALOMI extends Universal Manipulation Interface (UMI) with egocentric sensing to collect ego-view and wrist-view observations along with head-hand trajectories at scale. We further propose a manifold-constrained controller that plans in a learned latent behavior manifold to enable precise and robust head-hand tracking in the world frame. To bridge the human-to-humanoid gap, we perform ego-view alignment and introduce a controller-aware reference trajectory adaptation to reduce mismatch in both observation and action execution. We validate HALOMI on a Unitree G1 humanoid robot with an actuated neck across five real-world tasks involving navigation, grasping, bimanual manipulation, whole-body coordination, and dynamic behaviors. Across the three quantitatively evaluated tasks, HALOMI achieves an average success rate of 85\%, while additional qualitative demonstrations show its ability to support dynamic tossing and deep-squat grasping.

人形机器人示范学习主动感知运动操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。