arXiv:2603.28243cs.ROcs.SY2026-03被引 1

用参数化MPC逼近价值函数,让机器人走路更稳更快

Cost-Matching Model Predictive Control for Efficient Reinforcement Learning in Humanoid Locomotion

  • 用中心质量动力学的参数化MPC学习最优动作价值
  • 仿真中比人工调参基线提升运动性能与抗干扰能力
  • 适合需要高效强化学习的复杂机器人控制场景

本文提出一种基于模型预测控制(MPC)的强化学习框架,用于高效的人形机器人步态控制。采用带有质心动力学的参数化MPC,通过高保真闭环数据训练其逼近动作价值函数。具体地,沿记录的状态-动作轨迹评估MPC的代价至终值,并更新参数以最小化预测值与实测回报之间的差异。该方法实现高效的梯度学习,同时避免了训练过程中反复求解MPC带来的计算负担。在商用人形平台的仿真中验证了该方法,结果表明其相比手动调参基线,显著提升了运动性能与对模型失配及外部扰动的鲁棒性。

原文摘要 · Abstract (English)

In this paper, we propose a cost-matching approach for optimal humanoid locomotion within a Model Predictive Control (MPC)-based Reinforcement Learning (RL) framework. A parameterized MPC formulation with centroidal dynamics is trained to approximate the action-value function obtained from high-fidelity closed-loop data. Specifically, the MPC cost-to-go is evaluated along recorded state-action trajectories, and the parameters are updated to minimize the discrepancy between MPC-predicted values and measured returns. This formulation enables efficient gradient-based learning while avoiding the computational burden of repeatedly solving the MPC problem during training. The proposed method is validated in simulation using a commercial humanoid platform. Results demonstrate improved locomotion performance and robustness to model mismatch and external disturbances compared with manually tuned baselines.

强化学习机器人控制模型预测人形机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。