用教师策略指导学生策略,提升人形机器人在复杂环境中的行走稳定性。
Distillation-PPO: A Novel Two-Stage Reinforcement Learning Framework for Humanoid Robot Perceptive Locomotion
- 先在理想环境中训练教师策略,再通过知识蒸馏将经验转移给学生策略。
- 在仿真中训练效率更高,真实场景下鲁棒性更强,泛化能力显著提升。
- 适合需要稳定行走控制的机器人研发人员,尤其关注真实世界部署者。
近年来,人形机器人因其高度环境适应性和类人特征受到学术界和工业界的广泛关注。随着强化学习的发展,人形机器人行走控制已取得显著进展。然而,面对复杂环境与不规则地形时,现有方法仍面临挑战。目前感知运动控制主要分为两阶段方法和端到端方法。两阶段方法先在模拟环境中训练教师策略,再利用如DAgger等蒸馏技术,将所学的隐含特征或动作作为优势信息传递给学生策略;而端到端方法则直接在部分可观测马尔可夫决策过程(POMDP)中通过强化学习学习策略,但因缺乏教师监督,常导致训练困难且实际应用中性能不稳定。本文提出一种新型两阶段感知运动框架,结合在完全可观测马尔可夫决策过程(MDP)中训练的教师策略,对学生策略进行正则化与监督,同时利用强化学习特性确保学生策略能在POMDP中持续学习,从而提升模型上限。实验表明,该框架在仿真环境中具有更高的训练效率与稳定性,并在真实场景中展现出更强的鲁棒性与泛化能力。
原文摘要 · Abstract (English)
In recent years, humanoid robots have garnered significant attention from both academia and industry due to their high adaptability to environments and human-like characteristics. With the rapid advancement of reinforcement learning, substantial progress has been made in the walking control of humanoid robots. However, existing methods still face challenges when dealing with complex environments and irregular terrains. In the field of perceptive locomotion, existing approaches are generally divided into two-stage methods and end-to-end methods. Two-stage methods first train a teacher policy in a simulated environment and then use distillation techniques, such as DAgger, to transfer the privileged information learned as latent features or actions to the student policy. End-to-end methods, on the other hand, forgo the learning of privileged information and directly learn policies from a partially observable Markov decision process (POMDP) through reinforcement learning. However, due to the lack of supervision from a teacher policy, end-to-end methods often face difficulties in training and exhibit unstable performance in real-world applications. This paper proposes an innovative two-stage perceptive locomotion framework that combines the advantages of teacher policies learned in a fully observable Markov decision process (MDP) to regularize and supervise the student policy. At the same time, it leverages the characteristics of reinforcement learning to ensure that the student policy can continue to learn in a POMDP, thereby enhancing the model's upper bound. Our experimental results demonstrate that our two-stage training framework achieves higher training efficiency and stability in simulated environments, while also exhibiting better robustness and generalization capabilities in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。