通过模型假设正则化,让仿人机器人在复杂地形上稳定跑出1.5米/秒
PPF: Pre-training and Preservative Fine-tuning of Humanoid Locomotion via Model-Assumption-based Regularization
- 先模仿模型控制器预训练,再用强化学习微调,关键在状态可信时才对齐动作
- 实测在滑、坡、不平、沙地等复杂地形均实现1.5米/秒稳定行走
- 适合做仿人机器人运动控制或想避免灾难性遗忘的研究者
由于仿人机器人运动本身具有高维动态复杂性,并需适应多变不可预测的环境,该任务极具挑战。本文提出一种新型学习框架,有效训练仿人机器人运动策略,使其既模仿模型驱动控制器的行为,又能扩展至更复杂的运动任务,如复杂地形和高速度指令。框架包含三个关键部分:通过模仿模型控制器进行预训练,通过强化学习进行微调,以及在微调过程中引入基于模型假设的正则化(MAR)。其中,MAR仅在模型假设成立的状态下使策略与模型控制器动作对齐,从而防止灾难性遗忘。我们在全尺寸仿人机器人Digit上通过全面的仿真测试和硬件实验验证了该框架,实现了1.5米/秒的前进速度,并在滑、坡、不平、沙地等多种复杂地形上展现出鲁棒的运动能力。
原文摘要 · Abstract (English)
Humanoid locomotion is a challenging task due to its inherent complexity and high-dimensional dynamics, as well as the need to adapt to diverse and unpredictable environments. In this work, we introduce a novel learning framework for effectively training a humanoid locomotion policy that imitates the behavior of a model-based controller while extending its capabilities to handle more complex locomotion tasks, such as more challenging terrain and higher velocity commands. Our framework consists of three key components: pre-training through imitation of the model-based controller, fine-tuning via reinforcement learning, and model-assumption-based regularization (MAR) during fine-tuning. In particular, MAR aligns the policy with actions from the model-based controller only in states where the model assumption holds to prevent catastrophic forgetting. We evaluate the proposed framework through comprehensive simulation tests and hardware experiments on a full-size humanoid robot, Digit, demonstrating a forward speed of 1.5 m/s and robust locomotion across diverse terrains, including slippery, sloped, uneven, and sandy terrains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。