arXiv:2604.01830cs.LGcs.SY2026-04被引 1

用物理先验增强强化学习,高效实现电网拓扑控制。

Physics Informed Reinforcement Learning with Gibbs Priors for Topology Control in Power Grids

  • 结合吉布斯先验与图神经网络,压缩可行动作空间。
  • 在三个真实场景中均达近最优性能,最快提升6倍。
  • 适合需快速响应的电网安全调控场景。

电网拓扑控制是极具挑战的序列决策问题,因动作空间随电网规模呈组合爆炸增长,且通过仿真评估动作代价高昂。本文提出一种物理信息强化学习框架,融合半马尔可夫控制与吉布斯先验,编码系统物理规律于动作空间。仅当电网进入危险状态时才做决策,由图神经网络代理模型预测可行拓扑动作后的过载风险。基于预测结果构建物理信息吉布斯先验,既筛选出小规模状态相关候选集,又重加权策略对数概率以指导动作选择。该方法显著降低探索难度与在线仿真开销,同时保持学习策略灵活性。在三个难度递增的真实基准环境上评估:首个场景下性能接近理想基准,速度提升约6倍;第二个场景中达到94.6%的理想奖励,决策时间降低约200倍;最复杂场景中,相比PPO基线,奖励最高提升255%,存活步数提升284%,且仍比强工程基线快约2.5倍。结果表明该方法为电网拓扑控制提供了高效可靠的新机制。

原文摘要 · Abstract (English)

Topology control for power grid operation is a challenging sequential decision making problem because the action space grows combinatorially with the size of the grid and action evaluation through simulation is computationally expensive. We propose a physics-informed Reinforcement Learning framework that combines semi-Markov control with a Gibbs prior, that encodes the system's physics, over the action space. The decision is only taken when the grid enters a hazardous regime, while a graph neural network surrogate predicts the post action overload risk of feasible topology actions. These predictions are used to construct a physics-informed Gibbs prior that both selects a small state-dependent candidate set and reweights policy logits before action selection. In this way, our method reduces exploration difficulty and online simulation cost while preserving the flexibility of a learned policy. We evaluate the approach in three realistic benchmark environments of increasing difficulty. Across all settings, the proposed method achieves a strong balance between control quality and computational efficiency: it matches oracle-level performance while being approximately $6\times$ faster on the first benchmark, reaches $94.6\%$ of oracle reward with roughly $200\times$ lower decision time on the second one, and on the most challenging benchmark improves over a PPO baseline by up to $255\%$ in reward and $284\%$ in survived steps while remaining about $2.5\times$ faster than a strong specialized engineering baseline. These results show that our method provides an effective mechanism for topology control in power grids.

强化学习电网控制物理信息图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。