arXiv:2510.10455cs.ROcs.SY2025-10中稿 · while submitting t…

利用对称性设计奖励函数,让四足机器人自动生成多种步态并自由切换。

Towards Dynamic Quadrupedal Gaits: A Symmetry-Guided RL Hierarchy Enables Free Gait Transitions at Varying Speeds

  • 基于时间、形态和时间反演对称性设计奖励函数
  • 在不同速度下实现踏步模式的平滑过渡,无需预设轨迹
  • 适用于四足机器人硬件部署,减少调参工作量

四足机器人具有多种可行步态,但生成特定踏步序列通常需繁琐地人工调整触地与离地时刻及各腿的全向约束。本文提出一种统一的强化学习框架,通过利用动态足式系统内在的对称性与速度-周期关系,生成多样化的四足步态。我们设计了一种对称性引导的奖励函数,融合时间、形态和时间反演对称性。聚焦于保持的对称性和自然动力学,该方法无需预定义轨迹,可实现蹬步、跳跃、半跳跃和飞奔等多样运动模式间的平滑切换。在Unitree Go2机器人上实现,该方法在仿真与实机测试中均表现出跨速度范围的鲁棒性能,显著提升步态适应性,且无需大量奖励调参或显式足部定位控制。本工作为动态行走策略提供了新见解,凸显了对称性在机器人步态设计中的关键作用。

原文摘要 · Abstract (English)

Quadrupedal robots exhibit a wide range of viable gaits, but generating specific footfall sequences often requires laborious expert tuning of numerous variables, such as touch-down and lift-off events and holonomic constraints for each leg. This paper presents a unified reinforcement learning framework for generating versatile quadrupedal gaits by leveraging the intrinsic symmetries and velocity-period relationship of dynamic legged systems. We propose a symmetry-guided reward function design that incorporates temporal, morphological, and time-reversal symmetries. By focusing on preserved symmetries and natural dynamics, our approach eliminates the need for predefined trajectories, enabling smooth transitions between diverse locomotion patterns such as trotting, bounding, half-bounding, and galloping. Implemented on the Unitree Go2 robot, our method demonstrates robust performance across a range of speeds in both simulations and hardware tests, significantly improving gait adaptability without extensive reward tuning or explicit foot placement control. This work provides insights into dynamic locomotion strategies and underscores the crucial role of symmetries in robotic gait design.

四足机器人强化学习步态生成对称性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。