arXiv:2409.04467eess.SYcs.LG2024-09被引 4

通过数据驱动方法分解电网控制问题,提升强化学习效率。

State and Action Factorization in Power Grids

  • 基于数据自动识别状态与动作的相关性,分组形成独立子问题
  • 在Grid2Op基准上验证效果,与领域专家分析一致
  • 为分布式强化学习提供理论基础,适合电网优化研究者

为实现零排放目标,可再生能源占比不断提升,使电网控制面临更大挑战。近年来的Learning To Run a Power Network(L2RPN)系列竞赛推动了强化学习(RL)在辅助人工调度中的应用。现有解决方案普遍限制动作空间,或采用单一智能体全局控制,或多个独立智能体在变电站层级并行。本文提出一种无需领域先验的算法,完全基于数据估计状态与动作组件间的相关性。高度相关的状态-动作对被分组,形成更简单、可能独立的子问题,从而支持差异化的学习过程,降低计算与数据需求。该方法在使用Grid2Op模拟器构建的电网基准上进行了验证,结果与领域专家分析相符。基于此,本文为分布式强化学习在电网优化中的应用奠定了理论基础。

原文摘要 · Abstract (English)

The increase of renewable energy generation towards the zero-emission target is making the problem of controlling power grids more and more challenging. The recent series of competitions Learning To Run a Power Network (L2RPN) have encouraged the use of Reinforcement Learning (RL) for the assistance of human dispatchers in operating power grids. All the solutions proposed so far severely restrict the action space and are based on a single agent acting on the entire grid or multiple independent agents acting at the substations level. In this work, we propose a domain-agnostic algorithm that estimates correlations between state and action components entirely based on data. Highly correlated state-action pairs are grouped together to create simpler, possibly independent subproblems that can lead to distinct learning processes with less computational and data requirements. The algorithm is validated on a power grid benchmark obtained with the Grid2Op simulator that has been used throughout the aforementioned competitions, showing that our algorithm is in line with domain-expert analysis. Based on these results, we lay a theoretically-grounded foundation for using distributed reinforcement learning in order to improve the existing solutions.

强化学习电网控制分布式学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。