arXiv:2511.20977eess.SYcs.LG2025-11被引 1

提出分布式强化学习算法,平衡多微网系统的经济性与可靠性。

Independent policy gradient-based reinforcement learning for economic and reliable energy management of multi-microgrid systems

  • 各微网独立更新策略,通过均值-方差博弈建模经济与可靠性
  • 算法在已知/未知模型下均实现稳定优化,提升系统整体性能
  • 适合需要分布式协同的智能电网场景,尤其关注可靠性保障

效率与可靠性对集成间歇性、分布式可再生能源的多微网系统(MMSs)至关重要。本文研究分布式架构下的经济可靠能量管理问题,各微网独立以去中心化方式更新策略,协同优化长期系统性能。引入主网交互功率的均值与方差作为经济性与可靠性的指标,将问题建模为均值-方差团队随机博弈(MV-TSG),传统基于期望累积奖励的方法不适用于方差度量。为此,提出一种完全分布式的独立策略梯度算法,并提供严格的收敛性分析,适用于已知模型参数的情况。对于大规模未知模型参数场景,进一步发展基于独立策略梯度的深度强化学习算法,实现数据驱动的策略优化。两种场景的数值实验验证了方法的有效性。所提方法充分挖掘了多微网系统的分布式计算能力,实现了经济性与运行可靠性的良好平衡。

原文摘要 · Abstract (English)

Efficiency and reliability are both crucial for energy management, especially in multi-microgrid systems (MMSs) integrating intermittent and distributed renewable energy sources. This study investigates an economic and reliable energy management problem in MMSs under a distributed scheme, where each microgrid independently updates its energy management policy in a decentralized manner to optimize the long-term system performance collaboratively. We introduce the mean and variance of the exchange power between the MMS and the main grid as indicators for the economic performance and reliability of the system. Accordingly, we formulate the energy management problem as a mean-variance team stochastic game (MV-TSG), where conventional methods based on the maximization of expected cumulative rewards are unsuitable for variance metrics. To solve MV-TSGs, we propose a fully distributed independent policy gradient algorithm, with rigorous convergence analysis, for scenarios with known model parameters. For large-scale scenarios with unknown model parameters, we further develop a deep reinforcement learning algorithm based on independent policy gradients, enabling data-driven policy optimization. Numerical experiments in two scenarios validate the effectiveness of the proposed methods. Our approaches fully leverage the distributed computational capabilities of MMSs and achieve a well-balanced trade-off between economic performance and operational reliability.

能量管理强化学习多微网分布式优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。