用分层强化学习+运行时安全防护,让电网控制器更安全、更通用。
Hierarchical Reinforcement Learning with Runtime Safety Shielding for Power Grid Operation

- 高层策略做长期决策,底层防护实时过滤危险动作。
- 在压力测试下存活时间更长,线路负载峰值降低37%。
- 无需重新训练即可跨电网部署,适合电力系统工程师使用。
强化学习在电网拓扑控制与拥堵管理中展现出潜力,但实际应用受限于严格的安全要求、罕见扰动下的脆弱性以及对未知电网拓扑的泛化能力差。在关键基础设施中,灾难性故障不可容忍,基于学习的控制器必须遵守硬性物理约束。本文提出一种安全约束的分层控制框架,将长周期决策与实时可行性验证分离:高层强化学习策略生成抽象控制动作,底层确定性运行时安全防护通过快速前向仿真过滤不安全动作。安全作为运行时不变量,独立于策略质量或训练分布。在Grid2Op基准套件上评估,包括正常条件、强制线路断开压力测试,以及零样本部署到ICAPS 2021大规模输电网(无需重训练)。结果表明,传统强化学习在压力下易崩溃,仅安全方法过于保守。而本方法在生存时长、线路负载峰值和零样本泛化方面均表现更优。说明电网控制中的安全与泛化应通过架构设计实现,而非复杂奖励工程,为可部署的学习型控制器提供可行路径。
原文摘要 · Abstract (English)
Reinforcement learning has shown promise for automating power-grid operation tasks such as topology control and congestion management. However, its deployment in real-world power systems remains limited by strict safety requirements, brittleness under rare disturbances, and poor generalization to unseen grid topologies. In safety-critical infrastructure, catastrophic failures cannot be tolerated, and learning-based controllers must operate within hard physical constraints. This paper proposes a safety-constrained hierarchical control framework for power-grid operation that explicitly decouples long-horizon decision-making from real-time feasibility enforcement. A high-level reinforcement learning policy proposes abstract control actions, while a deterministic runtime safety shield filters unsafe actions using fast forward simulation. Safety is enforced as a runtime invariant, independent of policy quality or training distribution. The proposed framework is evaluated on the Grid2Op benchmark suite under nominal conditions, forced line-outage stress tests, and zero-shot deployment on the ICAPS 2021 large-scale transmission grid without retraining. Results show that flat reinforcement learning policies are brittle under stress, while safety-only methods are overly conservative. In contrast, the proposed hierarchical and safety-aware approach achieves longer episode survival, lower peak line loading, and robust zero-shot generalization to unseen grids. These results indicate that safety and generalization in power-grid control are best achieved through architectural design rather than increasingly complex reward engineering, providing a practical path toward deployable learning-based controllers for real-world energy systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。