用多目标贝叶斯优化安全高效调优模型预测控制的强化学习参数。
Intersection of Reinforcement Learning and Bayesian Optimization for Intelligent Control of Industrial Processes: A Safe MPC-based DPG using Multi-Objective BO
- 将MPC强化学习与多目标贝叶斯优化结合,利用梯度信息指导搜索。
- 在模型不完美时仍实现样本高效、稳定且高性能的控制学习。
- 适合对安全性与可解释性要求高的工业过程智能控制场景。
基于模型预测控制(MPC)的强化学习(RL)为深度神经网络(DNN)驱动的强化学习提供了结构化且可解释的替代方案,具有更低的计算复杂度和更高的透明性。然而,标准的MPC-RL方法常面临收敛慢、因参数化受限导致策略学习次优,以及在线适应中的安全问题。为此,我们提出一种新框架,将MPC-RL与多目标贝叶斯优化(MOBO)相结合。该框架利用兼容确定性策略梯度(CDPG)方法估算强化学习阶段代价及其梯度,并将其作为噪声观测输入到采用期望超体积改进(EHVI)获取函数的MOBO算法中。这一融合实现了对MPC参数的高效且安全的调优,即使在模型存在偏差的情况下也能提升闭环性能。数值实验表明,该方法能在模型不完美条件下实现样本高效、稳定且高性能的控制学习。
原文摘要 · Abstract (English)
Model Predictive Control (MPC)-based Reinforcement Learning (RL) offers a structured and interpretable alternative to Deep Neural Network (DNN)-based RL methods, with lower computational complexity and greater transparency. However, standard MPC-RL approaches often suffer from slow convergence, suboptimal policy learning due to limited parameterization, and safety issues during online adaptation. To address these challenges, we propose a novel framework that integrates MPC-RL with Multi-Objective Bayesian Optimization (MOBO). The proposed MPC-RL-MOBO utilizes noisy evaluations of the RL stage cost and its gradient, estimated via a Compatible Deterministic Policy Gradient (CDPG) approach, and incorporates them into a MOBO algorithm using the Expected Hypervolume Improvement (EHVI) acquisition function. This fusion enables efficient and safe tuning of the MPC parameters to achieve improved closed-loop performance, even under model imperfections. A numerical example demonstrates the effectiveness of the proposed approach in achieving sample-efficient, stable, and high-performance learning for control systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。