arXiv:2511.00153cs.RO2025-11被引 27

通过模仿人类第一视角动作,让机器人学会手眼协同操作。

EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations

  • 从人类第一视角数据中学习手部与头部协同运动轨迹
  • 引入记忆增强策略应对快速头部视角变化,提升泛化能力
  • 适合研究人形机器人具身智能与仿生操控的团队

从人类示范中进行模仿学习为机器人技能获取提供了有前景的路径,但第一视角人类数据因身体差异带来根本性挑战。在操作过程中,人类会主动协调头部与手部动作,持续调整视角,并采用预操作视觉聚焦策略定位目标物体。这些行为产生动态、任务驱动的头部运动,静态机器人感知系统无法复现,导致分布偏移,显著降低策略性能。我们提出EgoMI(第一视角操作接口)框架,捕捉操作任务中末端执行器与主动头部运动的同步轨迹,生成可适配兼容类人形机器人形态的数据。为应对快速且大范围的头部视角变化,引入记忆增强策略,选择性融合历史观测。在配备可动摄像头头的双臂机器人上评估表明,显式建模头部运动的策略始终优于基线方法。结果表明,结合头部运动建模的协同手眼学习能有效弥合类人形机器人与人类之间的具身鸿沟,实现鲁棒的模仿学习。

原文摘要 · Abstract (English)

Imitation learning from human demonstrations offers a promising approach for robot skill acquisition, but egocentric human data introduces fundamental challenges due to the embodiment gap. During manipulation, humans actively coordinate head and hand movements, continuously reposition their viewpoint and use pre-action visual fixation search strategies to locate relevant objects. These behaviors create dynamic, task-driven head motions that static robot sensing systems cannot replicate, leading to a significant distribution shift that degrades policy performance. We present EgoMI (Egocentric Manipulation Interface), a framework that captures synchronized end-effector and active head trajectories during manipulation tasks, resulting in data that can be retargeted to compatible semi-humanoid robot embodiments. To handle rapid and wide-spanning head viewpoint changes, we introduce a memory-augmented policy that selectively incorporates historical observations. We evaluate our approach on a bimanual robot equipped with an actuated camera head and find that policies with explicit head-motion modeling consistently outperform baseline methods. Results suggest that coordinated hand-eye learning with EgoMI effectively bridges the human-robot embodiment gap for robust imitation learning on semi-humanoid embodiments. Project page: https://egocentric-manipulation-interface.github.io

模仿学习手眼协同第一视角类人机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。