arXiv:2505.11164cs.RO2025-05被引 54

用多专家蒸馏+强化学习,让机器人在复杂地形上又快又稳地跑。

Parkour in the Wild: Learning a General and Extensible Agile Locomotion Policy Using Multi-expert Distillation and RL Fine-tuning

  • 先学多种地形专家技能,再蒸馏成统一基础策略
  • 在真实3D扫描地形上微调,实现跨环境强泛化
  • 仅用深度图输入,适合实际野外部署

腿式机器人在轮式机器人无法进入的复杂地形中具有优势,适用于搜救或太空探索。然而现有控制方法难以在多样化的非结构化环境中泛化。本文提出一种结合多专家蒸馏与强化学习(RL)微调的新框架,实现敏捷运动控制。首先训练针对不同地形的专家策略以获取专项运动能力;随后通过DAgger算法将这些策略蒸馏为统一的基础策略;再在更广泛的地形集(包括真实3D扫描数据)上使用RL进行微调,提升泛化性能。该框架支持通过重复微调适应新地形。所提策略仅依赖深度图像作为外部感知输入,可在多样化非结构化环境中实现鲁棒导航。实验表明,相比现有方法,该策略能更有效地融合多地形技能于单一控制器中。在ANYmal D机器人上的部署验证了其在复杂环境中兼具敏捷性与鲁棒性的能力,树立了腿式机器人运动控制的新基准。

原文摘要 · Abstract (English)

Legged robots are well-suited for navigating terrains inaccessible to wheeled robots, making them ideal for applications in search and rescue or space exploration. However, current control methods often struggle to generalize across diverse, unstructured environments. This paper introduces a novel framework for agile locomotion of legged robots by combining multi-expert distillation with reinforcement learning (RL) fine-tuning to achieve robust generalization. Initially, terrain-specific expert policies are trained to develop specialized locomotion skills. These policies are then distilled into a unified foundation policy via the DAgger algorithm. The distilled policy is subsequently fine-tuned using RL on a broader terrain set, including real-world 3D scans. The framework allows further adaptation to new terrains through repeated fine-tuning. The proposed policy leverages depth images as exteroceptive inputs, enabling robust navigation across diverse, unstructured terrains. Experimental results demonstrate significant performance improvements over existing methods in synthesizing multi-terrain skills into a single controller. Deployment on the ANYmal D robot validates the policy's ability to navigate complex environments with agility and robustness, setting a new benchmark for legged robot locomotion.

机器人运动强化学习多专家蒸馏泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。