用强化学习优化电网故障扩散的实时应对策略
Real-Time Cascade Mitigation in Power Systems Using Influence Graph Improved by Reinforcement Learning
- 将影响图转化为马尔可夫决策过程,结合强化学习做实时决策
- 在IEEE 14/118节点系统上验证,主动断开线路可显著降低连锁故障风险
- 设计保守策略避免恶化系统状态,适合电力调度实时防护场景
尽管现代电力系统具有高可靠性,但随着可再生能源渗透率提升,连锁故障风险日益增加。实时缓解连锁故障需在不确定性下快速做出复杂操作决策。本文将影响图扩展为马尔可夫决策过程(MDP)模型,用于电力传输系统中的实时连锁故障缓解,考虑发电、负荷及初始故障的不确定性。该MDP包含“不作为”动作以支持保守决策,并采用强化学习求解。提出一种基于策略梯度的学习算法,初始策略对应未采取缓解措施的情况,且能处理无效动作。通过精心设计奖励函数,所学策略可在不恶化系统状态的前提下采取保守行动。该方法在IEEE 14节点和IEEE 118节点系统上验证,结果表明主动断开某些线路可有效降低连锁故障传播风险,且部分线路在多场景中持续表现为关键缓解节点。
原文摘要 · Abstract (English)
Despite high reliability, modern power systems with growing renewable penetration face an increasing risk of cascading outages. Real-time cascade mitigation requires fast, complex operational decisions under uncertainty. In this work, we extend the influence graph into a Markov decision process model (MDP) for real-time mitigation of cascading outages in power transmission systems, accounting for uncertainties in generation, load, and initial contingencies. The MDP includes a do-nothing action to allow for conservative decision-making and is solved using reinforcement learning. We present a policy gradient learning algorithm initialized with a policy corresponding to the unmitigated case and designed to handle invalid actions. The proposed learning method converges faster than the conventional algorithm. Through careful reward design, we learn a policy that takes conservative actions without deteriorating system conditions. The model is validated on the IEEE 14-bus and IEEE 118-bus systems. The results show that proactive line disconnections can effectively reduce cascading risk, and certain lines consistently emerge as critical in mitigating cascade propagation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。