arXiv:2503.15082cs.ROcs.AI2025-03被引 15

用对抗蒸馏让机器人跑得既灵活又自然。

StyleLoco: Generative Adversarial Distillation for Natural Humanoid Robot Locomotion

  • 先用强化学习训练出敏捷的教师策略,再通过多判别器蒸馏人类动作数据。
  • 在仿真和真实机器人上实现高精度、自然流畅的多样化步态。
  • 适合想提升机器人运动自然性与稳定性的研究者和工程师。

类人机器人需在不同速度和地形下具备多样化的运动能力并保持自然姿态。现有方法面临根本矛盾:基于手工奖励的强化学习可实现敏捷运动但生成不自然步态,而基于动作捕捉数据的生成对抗模仿学习虽能产生自然动作却存在训练不稳定且灵活性受限的问题。由于专家策略与人类动作数据间存在本质差异,二者融合困难。为此,我们提出StyleLoco,一种两阶段框架,通过生成对抗蒸馏(GAD)弥合这一差距。首先利用强化学习训练一个教师策略以获得动态敏捷的运动;随后采用多判别器架构,多个判别器同时从教师策略和动作捕捉数据中提取技能。该方法有效结合了强化学习的敏捷性与人类动作的自然流畅性,同时缓解了对抗训练中的常见不稳定性。通过大量仿真与真实世界实验,验证了StyleLoco使类人机器人能够以专家级精度完成多样化运动任务,并保留人类运动的自然美感,成功实现不同运动类型间的风格迁移,且在广泛命令输入下保持稳定运动。

原文摘要 · Abstract (English)

Humanoid robots are anticipated to acquire a wide range of locomotion capabilities while ensuring natural movement across varying speeds and terrains. Existing methods encounter a fundamental dilemma in learning humanoid locomotion: reinforcement learning with handcrafted rewards can achieve agile locomotion but produces unnatural gaits, while Generative Adversarial Imitation Learning (GAIL) with motion capture data yields natural movements but suffers from unstable training processes and restricted agility. Integrating these approaches proves challenging due to the inherent heterogeneity between expert policies and human motion datasets. To address this, we introduce StyleLoco, a novel two-stage framework that bridges this gap through a Generative Adversarial Distillation (GAD) process. Our framework begins by training a teacher policy using reinforcement learning to achieve agile and dynamic locomotion. It then employs a multi-discriminator architecture, where distinct discriminators concurrently extract skills from both the teacher policy and motion capture data. This approach effectively combines the agility of reinforcement learning with the natural fluidity of human-like movements while mitigating the instability issues commonly associated with adversarial training. Through extensive simulation and real-world experiments, we demonstrate that StyleLoco enables humanoid robots to perform diverse locomotion tasks with the precision of expertly trained policies and the natural aesthetics of human motion, successfully transferring styles across different movement types while maintaining stable locomotion across a broad spectrum of command inputs.

类人机器人运动生成对抗蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。