让机器人像人一样自然走动,还能抗干扰。
HiLo: Learning Whole-Body Human-like Locomotion with Motion Tracking Controller
- 用动作追踪控制器+强化学习,分解全身控制任务
- 在真实机器人上实现自然敏捷的行走,抗外力扰动
- 无需微调即可快速适配新任务,适合实际部署
深度强化学习(RL)已成为开发类人机器人行走控制器的有力方法。尽管以往的RL控制器展现出稳健的行走能力,但其动作往往缺乏人机场景所需的自然与敏捷性。本文提出HiLo(基于动作追踪的人类式运动),一种有效框架,用于学习具备人类般运动特征的RL策略。主要挑战在于复杂的奖励设计与领域随机化。HiLo通过基于RL的动作追踪控制器,结合简单的领域随机化手段(如随机力注入和动作延迟)克服这些问题。在该框架下,全身控制问题可拆分为开环控制部分与剩余部分由RL策略处理。此外,采用分布值函数以稳定训练过程,提升在扰动动力学下的累积奖励估计。实验表明,使用HiLo训练的动作追踪控制器可在真实系统中实现自然、敏捷的人类式行走,并对外部扰动具有鲁棒性。同时,通过残差机制可无需微调即适应不同任务需求,实现快速调整。
原文摘要 · Abstract (English)
Deep Reinforcement Learning (RL) has emerged as a promising method to develop humanoid robot locomotion controllers. Despite the robust and stable locomotion demonstrated by previous RL controllers, their behavior often lacks the natural and agile motion patterns necessary for human-centric scenarios. In this work, we propose HiLo (human-like locomotion with motion tracking), an effective framework designed to learn RL policies that perform human-like locomotion. The primary challenges of human-like locomotion are complex reward engineering and domain randomization. HiLo overcomes these issues by developing an RL-based motion tracking controller and simple domain randomization through random force injection and action delay. Within the framework of HiLo, the whole-body control problem can be decomposed into two components: One part is solved using an open-loop control method, while the residual part is addressed with RL policies. A distributional value function is also implemented to stabilize the training process by improving the estimation of cumulative rewards under perturbed dynamics. Our experiments demonstrate that the motion tracking controller trained using HiLo can perform natural and agile human-like locomotion while exhibiting resilience to external disturbances in real-world systems. Furthermore, we show that the motion patterns of humanoid robots can be adapted through the residual mechanism without fine-tuning, allowing quick adjustments to task requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。