用混合动作与全身控制,让机器人在真实家庭环境高效完成复杂操作。
HoMeR: Learning In-the-Wild Mobile Manipulation via Hybrid Imitation and Whole-Body Control
- 结合全局动作与精细调节,通过快速运动学控制器实现移动机械臂协同
- 仅需每任务20次示范,在3个仿真和3个真实任务中达成79.17%成功率
- 兼容视觉语言模型,可泛化至新物体、新布局和杂乱场景,适合家庭服务机器人
我们提出HoMeR,一种结合全身控制与混合动作模式的模仿学习框架,用于在真实家庭环境中实现移动操作。其核心是一个基于运动学的快速全身控制器,将期望的末端执行器位姿映射为移动基座与机械臂的协调运动。在此简化后的末端执行器动作空间中,HoMeR学习在长距离移动时预测绝对位姿,而在精细操作时预测相对位姿,将底层协调交由控制器处理,专注学习高层任务决策。我们在配备7自由度机械臂的全向移动机械臂上部署了HoMeR,测试了3个仿真和3个真实家庭任务(如开柜子、扫垃圾、整理枕头)。结果显示,仅需每任务20次示范,HoMeR在所有任务中总体成功率达79.17%,比最优基线平均高出29.17%。此外,HoMeR兼容视觉语言模型,能利用其互联网规模先验,更好地泛化至新物体外观、新布局及杂乱场景。总体而言,HoMeR突破了传统桌面设置限制,为样本高效、可泛化的日常室内操作提供了可扩展路径。代码、视频及补充材料见:http://homer-manip.github.io
原文摘要 · Abstract (English)
We introduce HoMeR, an imitation learning framework for mobile manipulation that combines whole-body control with hybrid action modes that handle both long-range and fine-grained motion, enabling effective performance on realistic in-the-wild tasks. At its core is a fast, kinematics-based whole-body controller that maps desired end-effector poses to coordinated motion across the mobile base and arm. Within this reduced end-effector action space, HoMeR learns to switch between absolute pose predictions for long-range movement and relative pose predictions for fine-grained manipulation, offloading low-level coordination to the controller and focusing learning on task-level decisions. We deploy HoMeR on a holonomic mobile manipulator with a 7-DoF arm in a real home. We compare HoMeR to baselines without hybrid actions or whole-body control across 3 simulated and 3 real household tasks such as opening cabinets, sweeping trash, and rearranging pillows. Across tasks, HoMeR achieves an overall success rate of 79.17% using just 20 demonstrations per task, outperforming the next best baseline by 29.17 on average. HoMeR is also compatible with vision-language models and can leverage their internet-scale priors to better generalize to novel object appearances, layouts, and cluttered scenes. In summary, HoMeR moves beyond tabletop settings and demonstrates a scalable path toward sample-efficient, generalizable manipulation in everyday indoor spaces. Code, videos, and supplementary material are available at: http://homer-manip.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。