arXiv:2601.17428cs.RO2026-01被引 4

让机器人自适应学习复杂地形行走,实现高速稳定运动。

Scaling Rough Terrain Locomotion with Automatic Curriculum Reinforcement Learning

  • 基于学习进度自动调整任务难度,无需预设难易顺序。
  • 四足机器人在多种地形上实现2.5米/秒的高速稳定行走。
  • 适合需要大规模复杂环境训练的机器人研究者。

课程学习在机器人学习中已展现出显著效果,但在复杂多样的任务空间中仍面临挑战,因缺乏明确的难度结构,难以定义先前方法所需的难度排序。本文提出基于学习进度的自动课程强化学习框架(LP-ACRL),通过在线估计智能体学习进度并自适应调整任务采样分布,实现无需先验难度信息的自动课程生成。采用LP-ACRL训练的ANYmal D四足机器人,在包括楼梯、斜坡、碎石地和低摩擦平面等多种地形上,实现了2.5米/秒线速度与3.0弧度/秒角速度下的稳定高速运动;而此前方法通常仅限于平坦地形上的高速或复杂地形上的低速。实验表明,LP-ACRL具备强可扩展性与实际应用价值,为复杂广泛任务空间中的课程生成研究提供了可靠基线。

原文摘要 · Abstract (English)

Curriculum learning has demonstrated substantial effectiveness in robot learning. However, it still faces limitations when scaling to complex, wide-ranging task spaces. Such task spaces often lack a well-defined difficulty structure, making the difficulty ordering required by previous methods challenging to define. We propose a Learning Progress-based Automatic Curriculum Reinforcement Learning (LP-ACRL) framework, which estimates the agent's learning progress online and adaptively adjusts the task-sampling distribution, thereby enabling automatic curriculum generation without prior knowledge of the difficulty distribution over the task space. Policies trained with LP-ACRL enable the ANYmal D quadruped to achieve and maintain stable, high-speed locomotion at 2.5 m/s linear velocity and 3.0 rad/s angular velocity across diverse terrains, including stairs, slopes, gravel, and low-friction flat surfaces--whereas previous methods have generally been limited to high speeds on flat terrain or low speeds on complex terrain. Experimental results demonstrate that LP-ACRL exhibits strong scalability and real-world applicability, providing a robust baseline for future research on curriculum generation in complex, wide-ranging robotic learning task spaces.

机器人强化学习自动课程四足行走

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。