用动物运动先验让四足机器人在复杂地形上自然行走
Motion Priors Reimagined: Adapting Flat-Terrain Skills for Complex Quadruped Mobility
- 分层强化学习:先学平地动物步态,再学复杂地形修正
- 在崎岖地形上保持流畅步态,比基线模型更稳定
- 适合想提升机器人泛化能力的研究者与工程师
基于强化学习的运动模仿方法在演示数据上训练可高效学习自然、生动的动作,且无需繁琐奖励设计,但常难以泛化到新环境。本文提出一种分层强化学习框架:先在平坦地面预训练低层策略,模仿动物运动以建立运动先验;随后高层目标条件策略在此基础上学习残差修正,实现感知行走、局部避障和目标导向导航,适应多样且崎岖的地形。仿真结果表明,所学残差能有效适应逐步增加难度的不平整地形,同时保留运动先验带来的步态特性。相比未使用运动先验的基线模型,在相似奖励设置下,本方法在运动正则化方面表现更优。真实世界实验中,ANYmal-D 四足机器人验证了该策略在复杂地形中实现类动物运动的泛化能力,展现出平稳高效的移动与局部导航性能。
原文摘要 · Abstract (English)
Reinforcement learning (RL)-based motion imitation methods trained on demonstration data can effectively learn natural and expressive motions with minimal reward engineering but often struggle to generalize to novel environments. We address this by proposing a hierarchical RL framework in which a low-level policy is first pre-trained to imitate animal motions on flat ground, thereby establishing motion priors. A subsequent high-level, goal-conditioned policy then builds on these priors, learning residual corrections that enable perceptive locomotion, local obstacle avoidance, and goal-directed navigation across diverse and rugged terrains. Simulation experiments illustrate the effectiveness of learned residuals in adapting to progressively challenging uneven terrains while still preserving the locomotion characteristics provided by the motion priors. Furthermore, our results demonstrate improvements in motion regularization over baseline models trained without motion priors under similar reward setups. Real-world experiments with an ANYmal-D quadruped robot confirm our policy's capability to generalize animal-like locomotion skills to complex terrains, demonstrating smooth and efficient locomotion and local navigation performance amidst challenging terrains with obstacles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。