arXiv:2507.01243cs.ROcs.LG2025-07被引 1

用自演化先验让机器人学会在复杂地形上单脚跳跃。

Jump-Start Reinforcement Learning with Self-Evolving Priors for Extreme Monopedal Locomotion

  • 分阶段训练,通过自我迭代优化策略引导。
  • 首次实现四足机器人在不规则地形上稳定单脚跳跃。
  • 适合研究极端运动控制与强化学习的学者。

强化学习在四足机器人敏捷运动中展现出巨大潜力,但直接训练其应对单足跳跃任务中的极端欠驱动和极端地形双重挑战仍极具难度,主要源于早期探索不稳定及奖励信号不可靠。为此,我们提出JumpER(通过自演化先验实现跳跃启动强化学习)框架,将策略学习划分为多阶段渐进复杂度训练。通过迭代自举已有策略动态生成自演化先验,逐步优化指导信号,从而在无需外部专家先验或手工奖励设计的情况下稳定探索与策略优化。结合结构化的三阶段课程——逐级演化动作模态、观测空间与任务目标,JumpER首次使四足机器人在不可预测地形上实现稳健单脚跳跃。显著成果包括:成功跨越最大60 cm的宽缝、不规则间距台阶,以及间距在15–35 cm之间变化的踏石。该方法为应对极端欠驱动与极端地形下的运动任务提供了系统性且可扩展的解决方案。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has shown great potential in enabling quadruped robots to perform agile locomotion. However, directly training policies to simultaneously handle dual extreme challenges, i.e., extreme underactuation and extreme terrains, as in monopedal hopping tasks, remains highly challenging due to unstable early-stage interactions and unreliable reward feedback. To address this, we propose JumpER (jump-start reinforcement learning via self-evolving priors), an RL training framework that structures policy learning into multiple stages of increasing complexity. By dynamically generating self-evolving priors through iterative bootstrapping of previously learned policies, JumpER progressively refines and enhances guidance, thereby stabilizing exploration and policy optimization without relying on external expert priors or handcrafted reward shaping. Specifically, when integrated with a structured three-stage curriculum that incrementally evolves action modality, observation space, and task objective, JumpER enables quadruped robots to achieve robust monopedal hopping on unpredictable terrains for the first time. Remarkably, the resulting policy effectively handles challenging scenarios that traditional methods struggle to conquer, including wide gaps up to 60 cm, irregularly spaced stairs, and stepping stones with distances varying from 15 cm to 35 cm. JumpER thus provides a principled and scalable approach for addressing locomotion tasks under the dual challenges of extreme underactuation and extreme terrains.

强化学习机器人运动自演化单足跳跃

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。