无需人类示范,机器人自主完成搬运与导航。
AdaptManip: Learning Adaptive Whole-Body Object Lifting and Delivery with Online Recurrent State Estimation
- 用递归网络实时追踪物体状态,抗遮挡。
- 强化学习训练全身控制策略,成功率显著提升。
- 零样本部署,适合真实场景的自主任务。
本文提出自适应全身运动-操作框架AdaptManip,实现人形机器人在无外部示范或遥控数据的情况下,完全自主地完成集成导航、物体搬运与投送任务。该框架包含三个耦合模块:(1) 基于递归网络的物体状态估计算法,在视野受限和遮挡条件下实时追踪物体;(2) 全身基线策略结合残差操纵控制,实现稳定搬运;(3) 基于激光雷达的全局定位估计器,提供抗漂移的定位能力。所有模块均在仿真中通过强化学习训练,并实现零样本部署到真实硬件。实验表明,AdaptManip在适应性和整体成功率上显著优于基线方法,包括基于模仿学习的方法;即使在遮挡情况下,精准的物体状态估计也能提升操作性能。进一步验证了在真实环境中人形机器人完成自主导航、物体搬运与投送的全流程能力。
原文摘要 · Abstract (English)
This paper presents Adaptive Whole-body Loco-Manipulation, AdaptManip, a fully autonomous framework for humanoid robots to perform integrated navigation, object lifting, and delivery. Unlike prior imitation learning-based approaches that rely on human demonstrations and are often brittle to disturbances, AdaptManip aims to train a robust loco-manipulation policy via reinforcement learning without human demonstrations or teleoperation data. The proposed framework consists of three coupled components: (1) a recurrent object state estimator that tracks the manipulated object in real time under limited field-of-view and occlusions; (2) a whole-body base policy for robust locomotion with residual manipulation control for stable object lifting and delivery; and (3) a LiDAR-based robot global position estimator that provides drift-robust localization. All components are trained in simulation using reinforcement learning and deployed on real hardware in a zero-shot manner. Experimental results show that AdaptManip significantly outperforms baseline methods, including imitation learning-based approaches, in adaptability and overall success rate, while accurate object state estimation improves manipulation performance even under occlusion. We further demonstrate fully autonomous real-world navigation, object lifting, and delivery on a humanoid robot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。