从人类示范中学习全身移动操作,无需机器人参与数据采集。
HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations
- 用第一视角感知扩展人类演示,实现无机器人的可迁移数据收集。
- 通过跨体感策略设计,解决人类与机器人在观测和动作上的差异问题。
- 支持长时序、双手协同、全身协调的复杂移动操作任务。
我们提出全身移动操作接口(HoMMI),一个无需机器人参与的数据采集与策略学习框架,直接从人类示范中学习全身移动操作。通过在现有UMI界面中引入第一视角感知,捕捉移动操作所需的全局上下文信息,实现便携、无机器人依赖且可扩展的数据采集。然而,直接引入第一视角感知会显著扩大人类与机器人在观测和动作空间中的体感差距,导致策略迁移困难。为此,我们设计了一种跨体感的手眼策略,包含与体感无关的视觉表征、放松的头部动作表示,以及基于机器人物理约束的全身协调控制器,以实现手眼轨迹。该方法支持需要双手协作、全身协调、导航与主动感知的长时序移动操作任务。结果详见:https://hommi-robot.github.io
原文摘要 · Abstract (English)
We present Whole-Body Mobile Manipulation Interface (HoMMI), a data collection and policy learning framework that learns whole-body mobile manipulation directly from robot-free human demonstrations. We augment UMI interfaces with egocentric sensing to capture the global context required for mobile manipulation, enabling portable, robot-free, and scalable data collection. However, naively incorporating egocentric sensing introduces a larger human-to-robot embodiment gap in both observation and action spaces, making policy transfer difficult. We explicitly bridge this gap with a cross-embodiment hand-eye policy design, including an embodiment agnostic visual representation; a relaxed head action representation; and a whole-body controller that realizes hand-eye trajectories through coordinated whole-body motion under robot-specific physical constraints. Together, these enable long-horizon mobile manipulation tasks requiring bimanual and whole-body coordination, navigation, and active perception. Results are best viewed on: https://hommi-robot.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。