arXiv:2510.04868eess.SYcs.AI2025-10被引 4

用模型预测控制引导强化学习,提升电力平衡收益

Model Predictive Control-Guided Reinforcement Learning for Implicit Balancing

  • 结合MPC的预测能力与RL的快速推理优势
  • 相比独立使用RL或MPC,收益提升超50%
  • 适合电力市场实时决策与储能优化场景

在欧洲,追求利润的平衡责任方可实时调整其日前申报,以协助输电系统运营商维持供需平衡。利用模型预测控制(MPC)策略挖掘此类隐性平衡机会虽能捕捉套利空间,但难以准确反映欧洲不平衡市场中的价格形成机制,且计算成本高。无模型强化学习(RL)方法执行速度快,但需大量数据训练,通常依赖实时和历史数据做决策。本文提出一种由MPC引导的强化学习方法,融合两者优势:既能将预测信息有效融入决策过程(如MPC),又保持了RL的快速推理能力。该方法在2023年比利时平衡数据上评估了隐性平衡电池控制问题的表现。首先,从多个角度分析了当前最优的独立RL与MPC方法,凸显其各自优劣。结果表明,所提方法相比独立RL和MPC分别实现16.15%和54.36%的套利利润提升。

原文摘要 · Abstract (English)

In Europe, profit-seeking balance responsible parties can deviate in real time from their day-ahead nominations to assist transmission system operators in maintaining the supply-demand balance. Model predictive control (MPC) strategies to exploit these implicit balancing strategies capture arbitrage opportunities, but fail to accurately capture the price-formation process in the European imbalance markets and face high computational costs. Model-free reinforcement learning (RL) methods are fast to execute, but require data-intensive training and usually rely on real-time and historical data for decision-making. This paper proposes an MPC-guided RL method that combines the complementary strengths of both MPC and RL. The proposed method can effectively incorporate forecasts into the decision-making process (as in MPC), while maintaining the fast inference capability of RL. The performance of the proposed method is evaluated on the implicit balancing battery control problem using Belgian balancing data from 2023. First, we analyze the performance of the standalone state-of-the-art RL and MPC methods from various angles, to highlight their individual strengths and limitations. Next, we show an arbitrage profit benefit of the proposed MPC-guided RL method of 16.15% and 54.36%, compared to standalone RL and MPC.

强化学习电力系统平衡市场储能优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。