arXiv:2502.06676cs.RO2025-02被引 7

让四足机器人自动切换步态,实现平滑敏捷的运动与故障自恢复。

Discovery of skill switching criteria for learning agile quadruped locomotion

  • 通过奖励函数融入接触模式,分步学习多种步态策略。
  • 高层策略根据目标距离动态加权,实现自然的步态切换。
  • 实机验证可流畅切换小跑、跳跃等步态,且能即时应对突发故障。

本文提出一种分层学习与优化框架,可学习并实现协调的多技能运动。所学策略能在追踪任意位置目标时自动、自然地切换步态,并在发生意外时迅速恢复。框架包含深度强化学习与优化过程:首先将接触模式嵌入奖励函数,独立学习各类步态策略;随后训练高层策略生成各策略权重,用于目标追踪任务中的多技能组合。技能切换阈值由奖励函数定义,并随学习进程通过外层优化循环更新。在模拟的Unitree A1四足机器人上完成多项综合任务验证;真实场景中成功展示小跑、跳跃、奔跃及其自然过渡。相比离散切换导致无法实现奔跃的方案,本方法实现了所有敏捷步态,且过渡更平滑连续。

原文摘要 · Abstract (English)

This paper develops a hierarchical learning and optimization framework that can learn and achieve well-coordinated multi-skill locomotion. The learned multi-skill policy can switch between skills automatically and naturally in tracking arbitrarily positioned goals and recover from failures promptly. The proposed framework is composed of a deep reinforcement learning process and an optimization process. First, the contact pattern is incorporated into the reward terms for learning different types of gaits as separate policies without the need for any other references. Then, a higher level policy is learned to generate weights for individual policies to compose multi-skill locomotion in a goal-tracking task setting. Skills are automatically and naturally switched according to the distance to the goal. The proper distances for skill switching are incorporated in reward calculation for learning the high level policy and updated by an outer optimization loop as learning progresses. We first demonstrated successful multi-skill locomotion in comprehensive tasks on a simulated Unitree A1 quadruped robot. We also deployed the learned policy in the real world showcasing trotting, bounding, galloping, and their natural transitions as the goal position changes. Moreover, the learned policy can react to unexpected failures at any time, perform prompt recovery, and resume locomotion successfully. Compared to discrete switch between single skills which failed to transition to galloping in the real world, our proposed approach achieves all the learned agile skills, with smoother and more continuous skill transitions.

四足机器人步态切换强化学习运动控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。