arXiv:2509.21231cs.RO2025-09被引 3

用模型引导的强化学习,让机器人手臂稳定应对走路时的晃动。

SEEC: Stable End-Effector Control with Model-Enhanced Residual Learning for Humanoid Loco-Manipulation

  • 用物理模型生成干扰,训练上半身策略补偿下肢扰动。
  • 在多种仿真和真实机器人上表现更优,误差降低30%以上。
  • 无需重新训练即可适应新走路方式,适合复杂任务场景。

手臂末端稳定对人形机器人协同运动任务至关重要,但受双足结构高自由度和固有动态不稳定性影响,实现困难。以往基于模型的控制器依赖精确动力学建模,难以捕捉摩擦、间隙等实际因素,导致性能下降;而学习方法虽能通过探索和域随机化缓解这些问题,在真实场景中仍有潜力,但常因过拟合训练条件需重训全身体系,且难以适应未见场景。为此,本文提出一种新型稳定末端控制框架(SEEC),结合模型增强的残差学习,通过带扰动生成器的模型引导强化学习,使上半身策略能够精准补偿由下肢运动引起的扰动。该设计使上半身策略在无需额外训练的情况下,可稳定应对多种未知行走控制器。我们在不同仿真环境及真实机器人Booster T1上验证了该框架,实验表明其持续优于基线方法,并能稳健处理多样且具挑战性的协同运动任务。

原文摘要 · Abstract (English)

Arm end-effector stabilization is essential for humanoid loco-manipulation tasks, yet it remains challenging due to the high degrees of freedom and inherent dynamic instability of bipedal robot structures. Previous model-based controllers achieve precise end-effector control but rely on precise dynamics modeling and estimation, which often struggle to capture real-world factors (e.g., friction and backlash) and thus degrade in practice. On the other hand, learning-based methods can better mitigate these factors via exploration and domain randomization, and have shown potential in real-world use. However, they often overfit to training conditions, requiring retraining with the entire body, and still struggle to adapt to unseen scenarios. To address these challenges, we propose a novel stable end-effector control (SEEC) framework with model-enhanced residual learning that learns to achieve precise and robust end-effector compensation for lower-body induced disturbances through model-guided reinforcement learning (RL) with a perturbation generator. This design allows the upper-body policy to achieve accurate end-effector stabilization as well as adapt to unseen locomotion controllers with no additional training. We validate our framework in different simulators and transfer trained policies to the Booster T1 humanoid robot. Experiments demonstrate that our method consistently outperforms baselines and robustly handles diverse and demanding loco-manipulation tasks.

人形机器人末端控制强化学习协同运动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。