arXiv:2608.14114cs.LG2026-08中稿 · publication in ACM…被引 1

用强化学习优化电网拓扑,让电网在高风能接入下更稳定。

Learning to Run Power Networks: Effective AlphaZero-inspired Topological Control

  • 借鉴AlphaZero思想,用蒙特卡洛树搜索主动规划电网开关动作
  • 98.43%的最高存活率,远超传统PPO方法
  • 简单二元生存奖励比复杂目标函数更有效,适合电力系统工程师

随着可再生能源波动性增强,现代电网面临更大压力。本文研究基于模型的AlphaZero-inspired方法,利用蒙特卡洛树搜索(MCTS)实现电网拓扑的主动调控。相比传统重调度方式,拓扑操作成本更低、响应更快。然而,其应用受限于庞大的组合动作空间和严格的运行约束。我们系统评估了奖励函数、观测密度与搜索引导对智能体存活率的影响。结果表明,优化后的AlphaZero方法达到98.43%的峰值存活率,显著优于近端策略优化(PPO)变体。发现不依赖预训练策略或价值函数进行MCTS可提升训练效率;而简单的二元生存奖励比多目标复杂函数更具搜索引导效果。研究表明,尽管AlphaZero框架强大,但纯强化学习仍不足;高效可靠的系统需融合领域启发式规则、二元奖励及线路负荷的受限观测空间。

原文摘要 · Abstract (English)

As the integration of volatile renewable energy sources increases the strain on modern power grids, the use of Reinforcement Learning (RL) for autonomous topological reconfiguration has emerged as a promising research field to keep strained grids stable and operational. Compared to traditional redispatching measures, topological actions offer a cheaper and more cost-effective way to manage grid congestion. However, their implementation is hindered by a vast combinatorial action space and strict operational constraints. This paper investigates the effectiveness of model-based AlphaZero-inspired approaches that utilize Monte Carlo Tree Search (MCTS) for proactive grid management. We systematically evaluate how reward functions, observation density, and search guidance influence an agent's survivability. Our results demonstrate that the optimized AlphaZero approach achieves a peak survivability of 98.43%, significantly outperforming the proximal policy optimization (PPO) variant. We find that conducting the MCTS without guidance from a prior learned policy or value function can enhance training efficiency, and that a straightforward binary survival reward provides more effective search guidance than complex, multi-objective functions. Our findings demonstrate that while AlphaZero is a powerful framework for topological control, pure reinforcement learning is not sufficient; rather, an effective and reliable system requires a 'minimalist' integration of domain-specific heuristics, binary rewards, and a restricted observation space of line loads.

电网控制强化学习AlphaZero拓扑优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。