arXiv:2509.02727cs.RO2025-09被引 1

用强化学习让四足机器人学会攀爬跳跃,效果媲美专业模型但训练成本更低。

Acrobotics: A Generalist Approach to Quadrupedal Robots' Parkour

  • 采用通用强化学习算法,通过试错训练四足机器人动态运动能力。
  • 仅用25%的训练智能体数量,性能接近顶尖专家混合模型。
  • 揭示了通用策略成功的关键组件和核心影响因素。

相较于轮式机器人,四足机器人在攀爬、蹲伏、跨越障碍和上楼梯等方面具有显著优势,更适用于复杂非结构化地形。然而,这些动作需要精确的时间协调与复杂的智能体-环境交互,且腿式运动易发生打滑或绊倒,传统建模方法难以设计出鲁棒控制器。相比之下,强化学习可通过试错实现最优控制。本文提出一种通用强化学习算法,用于四足机器人在动态运动场景中的行为学习。所学策略在性能上可媲美采用专家混合方法训练的顶尖专用策略,且训练时仅需其25%的智能体数量。实验还揭示了通用运动策略的关键构成及成功的主要因素。

原文摘要 · Abstract (English)

Climbing, crouching, bridging gaps, and walking up stairs are just a few of the advantages that quadruped robots have over wheeled robots, making them more suitable for navigating rough and unstructured terrain. However, executing such manoeuvres requires precise temporal coordination and complex agent-environment interactions. Moreover, legged locomotion is inherently more prone to slippage and tripping, and the classical approach of modeling such cases to design a robust controller thus quickly becomes impractical. In contrast, reinforcement learning offers a compelling solution by enabling optimal control through trial and error. We present a generalist reinforcement learning algorithm for quadrupedal agents in dynamic motion scenarios. The learned policy rivals state-of-the-art specialist policies trained using a mixture of experts approach, while using only 25% as many agents during training. Our experiments also highlight the key components of the generalist locomotion policy and the primary factors contributing to its success.

四足机器人强化学习通用智能体运动控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。