用多目标强化学习优化电网拓扑,平衡拥堵、操作频率与稳定性。
Multi-Objective Reinforcement Learning for Power Grid Topology Control
- 采用DOL与MOPPO算法生成多目标最优调控策略。
- 在故障情况下减少30%电网崩溃风险,训练效率提升20%。
- 适合电网调度员理解不同目标间的权衡,指导实际运行决策。
随着各领域电气化程度提高,输电网络拥堵问题日益严重。通过变电站重构实现的拓扑控制虽可缓解拥堵,但其潜力尚未在实际运营中充分挖掘。核心挑战在于如何建模该问题以契合运营商的多重目标与约束。本文研究将多目标强化学习(MORL)应用于电网拓扑控制,提出基于深度乐观线性支持(DOL)和多目标近端策略优化(MOPPO)的方法,生成一组帕累托最优策略,平衡线路负载、拓扑偏差与开关频率等冲突目标。初步案例研究显示,该方法能有效揭示目标间权衡,并优于随机搜索基线的帕累托前沿逼近能力。相较于常见单目标强化学习策略,在故障情景下预防电网失稳的成功率高出30%,且在训练预算缩减时仍保持20%的性能优势。
原文摘要 · Abstract (English)
Transmission grid congestion increases as the electrification of various sectors requires transmitting more power. Topology control, through substation reconfiguration, can reduce congestion but its potential remains under-exploited in operations. A challenge is modeling the topology control problem to align well with the objectives and constraints of operators. Addressing this challenge, this paper investigates the application of multi-objective reinforcement learning (MORL) to integrate multiple conflicting objectives for power grid topology control. We develop a MORL approach using deep optimistic linear support (DOL) and multi-objective proximal policy optimization (MOPPO) to generate a set of Pareto-optimal policies that balance objectives such as minimizing line loading, topological deviation, and switching frequency. Initial case studies show that the MORL approach can provide valuable insights into objective trade-offs and improve Pareto front approximation compared to a random search baseline. The generated multi-objective RL policies are 30% more successful in preventing grid failure under contingencies and 20% more effective when training budget is reduced - compared to the common single objective RL policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。