用晃动桥训练四足机器人,提升其在振动地面的行走稳定性。
Bridge the Gap: Enhancing Quadruped Locomotion with Vertical Ground Perturbations
- 在模拟中用2.0Hz振动摇桥训练机器人,结合强化学习策略。
- 实测显示,晃桥训练的机器人比平地训练更稳定,可零样本迁移至真实场景。
- 适合研究机器人动态环境适应或强化学习应用的开发者参考。
四足机器人虽擅长复杂地形,但在垂直地面扰动(如振动表面)下的表现仍不充分。本研究通过在13.24米长的钢混结构晃桥(固有频率2.0 Hz)上训练Unitree Go2机器人,提升其鲁棒性。采用MuJoCo仿真环境,基于近端策略优化(PPO)算法,训练了15种不同步态(五种: trot、pace、bound、free、default)与三种训练条件(刚性桥及两种不同高度调节策略的晃桥)组合的运动策略。通过域随机化实现零样本迁移至真实桥面。结果表明,晃桥训练的策略在稳定性与适应性上显著优于刚性桥训练。该框架使机器人无需事先接触桥梁即可生成稳健步态,验证了基于仿真的强化学习在应对动态地面扰动中的潜力,为设计穿越振动环境的机器人提供新思路。
原文摘要 · Abstract (English)
Legged robots, particularly quadrupeds, excel at navigating rough terrains, yet their performance under vertical ground perturbations, such as those from oscillating surfaces, remains underexplored. This study introduces a novel approach to enhance quadruped locomotion robustness by training the Unitree Go2 robot on an oscillating bridge - a 13.24-meter steel-and-concrete structure with a 2.0 Hz eigenfrequency designed to perturb locomotion. Using Reinforcement Learning (RL) with the Proximal Policy Optimization (PPO) algorithm in a MuJoCo simulation, we trained 15 distinct locomotion policies, combining five gaits (trot, pace, bound, free, default) with three training conditions: rigid bridge and two oscillating bridge setups with differing height regulation strategies (relative to bridge surface or ground). Domain randomization ensured zero-shot transfer to the real-world bridge. Our results demonstrate that policies trained on the oscillating bridge exhibit superior stability and adaptability compared to those trained on rigid surfaces. Our framework enables robust gait patterns even without prior bridge exposure. These findings highlight the potential of simulation-based RL to improve quadruped locomotion during dynamic ground perturbations, offering insights for designing robots capable of traversing vibrating environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。