融合强化学习与模型预测控制,实现四足机器人实时步态自适应调整。
Real-Time Gait Adaptation for Quadrupeds using Model Predictive Control and Reinforcement Learning
- 用MPPI优化动作与步态变量,结合梦境模型奖励函数。
- 能耗降低36.48%,速度跟踪精准,步态平滑切换。
- 适合需动态环境适应的四足机器人运动控制研究者。
无模型强化学习(RL)已实现四足机器人的灵活运动;但策略常收敛于单一步态,导致性能欠佳。传统模型预测控制(MPC)虽能获得特定任务最优策略,却难以适应复杂环境变化。为此,我们提出一种基于模型预测路径积分(MPPI)与Dreamer模块的优化框架,实现在连续步态空间中的实时步态自适应。每一步中,MPPI联合优化动作与步态变量,利用学习到的Dreamer奖励函数,兼顾速度追踪、能量效率、稳定性及平滑过渡,同时惩罚突变步态。引入学习的价值函数作为终值奖励,使规划扩展至无限时域。在Unitree Go1仿真平台上评估,结果表明目标速度变化时平均能耗降低达36.48%,同时保持高精度速度跟踪与任务适配的自适应步态。
原文摘要 · Abstract (English)
Model-free reinforcement learning (RL) has enabled adaptable and agile quadruped locomotion; however, policies often converge to a single gait, leading to suboptimal performance. Traditionally, Model Predictive Control (MPC) has been extensively used to obtain task-specific optimal policies but lacks the ability to adapt to varying environments. To address these limitations, we propose an optimization framework for real-time gait adaptation in a continuous gait space, combining the Model Predictive Path Integral (MPPI) algorithm with a Dreamer module to produce adaptive and optimal policies for quadruped locomotion. At each time step, MPPI jointly optimizes the actions and gait variables using a learned Dreamer reward that promotes velocity tracking, energy efficiency, stability, and smooth transitions, while penalizing abrupt gait changes. A learned value function is incorporated as terminal reward, extending the formulation to an infinite-horizon planner. We evaluate our framework in simulation on the Unitree Go1, demonstrating an average reduction of up to 36.48 % in energy consumption across varying target speeds, while maintaining accurate tracking and adaptive, task-appropriate gaits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。