用强化学习优化微电网控制,提速同时保持高精度。
Integrating Reinforcement Learning and Model Predictive Control with Applications to Microgrids
- 用强化学习决定离散动作,让MPC从混合整数规划变线性规划
- 仿真显示计算时间大幅降低,可行性高且近似最优
- 适合需要实时控制的能源系统研究者
本文提出一种融合强化学习与模型预测控制(MPC)的方法,高效解决混合逻辑动态系统中的有限时域最优控制问题。这类系统的优化控制需在线求解混合整数线性规划(MILP),面临维度灾难。本方法通过将离散变量与连续变量决策解耦,由强化学习确定离散变量,使MPC的在线优化问题从混合整数线性规划简化为线性规划,显著降低计算时间。核心贡献是定义了分拆的Q函数,使在组合动作空间中的学习问题可处理。我们采用循环神经网络近似该分拆Q函数,并展示其在强化学习中的应用。基于真实数据的微电网系统仿真表明,该方法大幅减少MPC在线计算时间,同时保持高可行性与低次优性。
原文摘要 · Abstract (English)
This work proposes an approach that integrates reinforcement learning and model predictive control (MPC) to solve finite-horizon optimal control problems in mixed-logical dynamical systems efficiently. Optimization-based control of such systems with discrete and continuous decision variables entails the online solution of mixed-integer linear programs, which suffer from the curse of dimensionality. Our approach aims to mitigate this issue by decoupling the decision on the discrete variables from the decision on the continuous variables. In the proposed approach, reinforcement learning determines the discrete decision variables and simplifies the online optimization problem of the MPC controller from a mixed-integer linear program to a linear program, significantly reducing the computational time. A fundamental contribution of this work is the definition of the decoupled Q-function, which plays a crucial role in making the learning problem tractable in a combinatorial action space. We motivate the use of recurrent neural networks to approximate the decoupled Q-function and show how they can be employed in a reinforcement learning setting. Simulation experiments on a microgrid system using real-world data demonstrate that the proposed method substantially reduces the online computation time of MPC while maintaining high feasibility and low suboptimality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。