arXiv:2503.18787cs.LGmath.OC2025-03被引 2

用柯尔莫哥洛夫模型提升非线性预测控制的样本效率

Sample-Efficient Reinforcement Learning of Koopman eNMPC

  • 将柯尔莫哥洛夫模型转化为可微策略,结合模型驱动强化学习
  • 在连续搅拌釜反应器上实现更优控制效果与更高采样效率
  • 融合物理先验知识进一步减少训练数据需求,适合工业控制场景

强化学习(RL)可用于通过优化动态模型或策略目标函数中的参数来调优数据驱动的(经济)非线性模型预测控制器((e)NMPC),以实现特定控制任务的最优性能。然而,样本效率至关重要。为此,本文结合模型驱动的强化学习算法与我们此前提出的将柯尔莫哥洛夫(e)NMPC转化为自动微分策略的方法。我们在文献中一个连续搅拌釜反应器(CSTR)的(e)NMPC案例研究中应用该方法。结果表明,该方法优于基准方法——即基于系统辨识构建的、未经进一步强化学习调优的数据驱动(e)NMPC,以及使用模型驱动强化学习训练的神经网络控制器,在控制性能和样本效率方面均表现更优。此外,通过引入关于系统动力学的部分先验知识进行物理信息学习,可进一步提高样本效率。

原文摘要 · Abstract (English)

Reinforcement learning (RL) can be used to tune data-driven (economic) nonlinear model predictive controllers ((e)NMPCs) for optimal performance in a specific control task by optimizing the dynamic model or parameters in the policy's objective function or constraints, such as state bounds. However, the sample efficiency of RL is crucial, and to improve it, we combine a model-based RL algorithm with our published method that turns Koopman (e)NMPCs into automatically differentiable policies. We apply our approach to an eNMPC case study of a continuous stirred-tank reactor (CSTR) model from the literature. The approach outperforms benchmark methods, i.e., data-driven eNMPCs using models based on system identification without further RL tuning of the resulting policy, and neural network controllers trained with model-based RL, by achieving superior control performance and higher sample efficiency. Furthermore, utilizing partial prior knowledge about the system dynamics via physics-informed learning further increases sample efficiency.

强化学习模型预测控制样本效率柯尔莫哥洛夫

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。