arXiv:2501.16590cs.ROcs.SY2025-01被引 11

对比强化学习与模型预测控制在四足机器人行走中的表现。

Benchmarking Model Predictive Control and Reinforcement Learning Based Control for Legged Robot Locomotion in MuJoCo Simulation

  • 在MuJoCo中对Go1机器人的直行任务进行对比测试
  • 强化学习抗干扰强且省电,但新地形适应差
  • 模型预测控制抗大扰动能力更强,关节控制更均衡

模型预测控制(MPC)和强化学习(RL)是两类主流的四足机器人控制方法。本文在MuJoCo仿真环境中,针对Unitree Go1四足机器人在恒定速度直行任务上,对两者进行标准化对比。评估指标包括抗扰性、能耗效率和地形适应性。结果表明:强化学习在抗干扰和节能方面表现优异,但其策略依赖特定环境训练,泛化能力弱;而模型预测控制凭借优化求解机制,在应对较大扰动时恢复能力更强,能实现关节间控制力的均衡分配。研究揭示了两种方法的优劣,为实际应用中控制策略选择提供依据。

原文摘要 · Abstract (English)

Model Predictive Control (MPC) and Reinforcement Learning (RL) are two prominent strategies for controlling legged robots, each with unique strengths. RL learns control policies through system interaction, adapting to various scenarios, whereas MPC relies on a predefined mathematical model to solve optimization problems in real-time. Despite their widespread use, there is a lack of direct comparative analysis under standardized conditions. This work addresses this gap by benchmarking MPC and RL controllers on a Unitree Go1 quadruped robot within the MuJoCo simulation environment, focusing on a standardized task-straight walking at a constant velocity. Performance is evaluated based on disturbance rejection, energy efficiency, and terrain adaptability. The results show that RL excels in handling disturbances and maintaining energy efficiency but struggles with generalization to new terrains due to its dependence on learned policies tailored to specific environments. In contrast, MPC shows enhanced recovery capabilities from larger perturbations by leveraging its optimization-based approach, allowing for a balanced distribution of control efforts across the robot's joints. The results provide a clear understanding of the advantages and limitations of both RL and MPC, offering insights into selecting an appropriate control strategy for legged robotic applications.

四足机器人强化学习模型预测控制仿真测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。