arXiv:2506.20036cs.ROcs.AI2025-06被引 2

分层强化学习让四足机器人更稳地跑过复杂地形

Hierarchical Reinforcement Learning and Value Optimization for Challenging Quadruped Locomotion

  • 高层策略在线优化低层策略的值函数,无需额外训练
  • 在多种地形上奖励更高、碰撞更少,包括训练时未见的难题地形
  • 适合做机器人运动控制研究或实际部署的工程师参考

我们提出一种新型分层强化学习框架,用于四足机器人在复杂地形上的运动控制。该方法采用双层结构:高层策略(HLP)为低层策略(LLP)选择最优步态目标,而低层策略基于在线策略的演员-评论家算法进行训练,并以脚点位置作为目标。我们设计的高层策略不依赖额外训练或环境采样,而是通过在线优化低层策略的已学价值函数实现。实验表明,相比端到端强化学习方法,该框架在多种地形上均能获得更高奖励、减少碰撞,且对训练中未遇过的更难地形也表现出更强适应能力。

原文摘要 · Abstract (English)

We propose a novel hierarchical reinforcement learning framework for quadruped locomotion over challenging terrain. Our approach incorporates a two-layer hierarchy in which a high-level policy (HLP) selects optimal goals for a low-level policy (LLP). The LLP is trained using an on-policy actor-critic RL algorithm and is given footstep placements as goals. We propose an HLP that does not require any additional training or environment samples and instead operates via an online optimization process over the learned value function of the LLP. We demonstrate the benefits of this framework by comparing it with an end-to-end reinforcement learning (RL) approach. We observe improvements in its ability to achieve higher rewards with fewer collisions across an array of different terrains, including terrains more difficult than any encountered during training.

强化学习机器人运动分层决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。