用模型监督强化学习,让机器人更稳地适应真实环境行走
Efficiently Learning Robust Torque-based Locomotion Through Reinforcement with Model-Based Supervision
- 用模型生成基础步态,再用强化学习学修正动作
- 在多种随机扰动下仍能稳定行走,泛化能力更强
- 适合做仿人机器人步态控制,尤其关注真实世界适应性
我们提出一种控制框架,将基于模型的双足步行与残差强化学习(RL)结合,以应对现实世界中的不确定性。该方法采用基于模型的控制器,包括发散运动分量(DCM)轨迹规划器和全身控制器,作为可靠的基准策略。为应对动态建模不准确和传感器噪声等问题,引入通过领域随机化训练的残差策略。关键在于,使用具有真值动力学访问权限的模型基虚拟策略,在训练中通过一种新型监督损失来指导残差策略。这种监督使策略能高效学习补偿未建模效应的修正行为,而无需大量奖励设计。实验表明,该方法在多种随机条件下均表现出更强的鲁棒性和泛化能力,为双足步行的仿真到现实迁移提供了可扩展的解决方案。
原文摘要 · Abstract (English)
We propose a control framework that integrates model-based bipedal locomotion with residual reinforcement learning (RL) to achieve robust and adaptive walking in the presence of real-world uncertainties. Our approach leverages a model-based controller, comprising a Divergent Component of Motion (DCM) trajectory planner and a whole-body controller, as a reliable base policy. To address the uncertainties of inaccurate dynamics modeling and sensor noise, we introduce a residual policy trained through RL with domain randomization. Crucially, we employ a model-based oracle policy, which has privileged access to ground-truth dynamics during training, to supervise the residual policy via a novel supervised loss. This supervision enables the policy to efficiently learn corrective behaviors that compensate for unmodeled effects without extensive reward shaping. Our method demonstrates improved robustness and generalization across a range of randomized conditions, offering a scalable solution for sim-to-real transfer in bipedal locomotion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。