arXiv:2511.09091cs.RO2025-11被引 1

用专家动作引导强化学习,让机器人更高效地学会自然行走。

APEX: Action Priors Enable Efficient Exploration for Robust Motion Tracking on Legged Robots

  • 将专家示范转化为随时间衰减的动作先验,引导探索
  • 无需部署时参考数据,提升样本效率与鲁棒性
  • 适合希望减少调参、实现多地形通用运动的开发者

从示范中学习自然、类动物般的运动已成为足式机器人领域的核心范式。尽管运动追踪技术取得进展,但现有方法仍需大量调参且依赖部署时的参考数据,限制了适应能力。本文提出APEX(Action Priors enable Efficient Exploration),一种可即插即用的运动追踪算法扩展,完全消除部署时对参考数据的依赖,提升样本效率并减少参数调优工作量。APEX通过引入衰减动作先验,将专家示范直接嵌入强化学习中:初期偏向专家动作以引导探索,随后逐步放开让策略自主探索。结合多评论家框架,平衡任务表现与运动风格。此外,单个策略可学习多种运动模式,并在不同地形和速度下迁移参考风格,同时对奖励设计变化保持鲁棒。我们在仿真环境及Unitree Go2机器人上进行了广泛验证。通过利用示范引导训练期探索而不施加显式偏好,APEX使足式机器人以更高稳定性、效率与泛化能力学习运动技能。我们相信该方法为指导驱动的强化学习在各类机器人任务(如运动与操作)中的自然技能获取开辟了新路径。代码与网页:https://marmotlab.github.io/APEX/

原文摘要 · Abstract (English)

Learning natural, animal-like locomotion from demonstrations has become a core paradigm in legged robotics. Despite the recent advancements in motion tracking, most existing methods demand extensive tuning and rely on reference data during deployment, limiting adaptability. We present APEX (Action Priors enable Efficient Exploration), a plug-and-play extension to state-of-the-art motion tracking algorithms that eliminates any dependence on reference data during deployment, improves sample efficiency, and reduces parameter tuning effort. APEX integrates expert demonstrations directly into reinforcement learning (RL) by incorporating decaying action priors, which initially bias exploration toward expert demonstrations but gradually allow the policy to explore independently. This is combined with a multi-critic framework that balances task performance with motion style. Moreover, APEX enables a single policy to learn diverse motions and transfer reference-like styles across different terrains and velocities, while remaining robust to variations in reward design. We validate the effectiveness of our method through extensive experiments in both simulation and on a Unitree Go2 robot. By leveraging demonstrations to guide exploration during RL training, without imposing explicit bias toward them, APEX enables legged robots to learn with greater stability, efficiency, and generalization. We believe this approach paves the way for guidance-driven RL to boost natural skill acquisition in a wide array of robotic tasks, from locomotion to manipulation. Website and code: https://marmotlab.github.io/APEX/.

强化学习足式机器人运动控制示范学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。