arXiv:2502.04925eess.SYcs.RO2025-02被引 1

用深度期望Sarsa与非线性时序差分学习,让非线性MPC自适应优化参数并稳定收敛。

Convergent NMPC-based Reinforcement Learning Using Deep Expected Sarsa and Nonlinear Temporal Difference Learning

  • 将NMPC作为动作价值函数,用神经网络逼近后续价值,输入包含学习参数以稳定训练。
  • 相比传统方法,实时计算量减半,闭环性能不变,且在仿真中实现局部最优收敛。
  • 适合需要高稳定性与实时性的复杂系统控制,如机器人、自动驾驶等场景。

本文提出一种基于强化学习的非线性模型预测控制(NMPC)方法,通过两种新策略学习NMPC方案的最优权重。首先,将控制器作为深度期望Sarsa中的当前动作价值函数,利用神经网络(NN)近似通常由次级NMPC获得的后续动作价值函数;同时将学习到的NMPC参数值加入神经网络输入,使网络能同时逼近动作价值并稳定学习过程。该方法在不降低闭环性能的前提下,使实时计算负担约减少一半。其次,结合梯度时序差分方法与参数化NMPC作为期望Sarsa的函数逼近器,解决非线性函数逼近下参数发散与不稳定的潜在问题。仿真结果表明,所提方法可收敛至局部最优解,且无发散或不稳定现象。

原文摘要 · Abstract (English)

In this paper, we present a learning-based nonlinear model predictive controller (NMPC) using an original reinforcement learning (RL) method to learn the optimal weights of the NMPC scheme, for which two methods are proposed. Firstly, the controller is used as the current action-value function of a deep Expected Sarsa where the subsequent action-value function, usually obtained with a secondary NMPC, is approximated with a neural network (NN). With respect to existing methods, we add to the NN's input the current value of the NMPC's learned parameters so that the network is able to approximate the action-value function and stabilize the learning performance. Additionally, with the use of the NN, the real-time computational burden is approximately halved without affecting the closed-loop performance. Secondly, we combine gradient temporal difference methods with a parametrized NMPC as a function approximator of the Expected Sarsa RL method to overcome the potential parameters' divergence and instability issues when nonlinearities are present in the function approximation. The simulation results show that the proposed approach converges to a locally optimal solution without instability problems.

强化学习模型预测控制非线性控制神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。