用异构智能体强化学习,高效协调微电网恢复供电。
Heterogeneous Multi-Agent Proximal Policy Optimization for Power Distribution System Restoration
- 不同微网设独立智能体,分层训练提升协同效率
- 在123节点和8500节点系统上恢复超95%负荷,收敛稳定
- 适合需实时恢复的复杂配电网,支持高并发场景
大范围停电后,配电网恢复需按序操作开关、重构馈线拓扑并协调分布式能源资源,受功率平衡、电压与热载荷等非线性约束限制,传统优化与基于价值的强化学习方法难以扩展。本文提出异构智能体近端策略优化(HAPPO),在异构代理强化学习框架下实现多微电网协同恢复。每个代理控制具有不同负载、分布式能源容量和开关数量的微电网。采用集中式评判者指导分散式执行者,确保在线策略学习稳定;通过融合物理规律的OpenDSS环境保障电气可行性。在IEEE 123节点与8500节点馈线系统上的实验表明,HAPPO在恢复电量、收敛稳定性及多种子可复现性上均优于PPO、QMIX、平均场强化学习等基线方法。在2400 kW发电容量限制下,两系统均恢复超过95%可用负荷,且执行延迟低,支持实际配电网的实时恢复应用。
原文摘要 · Abstract (English)
Restoring power distribution systems (PDSs) after large-scale outages requires sequential switching actions that reconfigure feeder topology and coordinate distributed energy resources (DERs) under nonlinear constraints, including power balance, voltage limits, and thermal ratings. These challenges limit the scalability of conventional optimization and value-based reinforcement learning (RL) approaches. This paper applies a Heterogeneous-Agent Reinforcement Learning (HARL) framework via Heterogeneous-Agent Proximal Policy Optimization (HAPPO) to enable coordinated restoration across interconnected microgrids. Each agent controls a distinct microgrid with different loads, DER capacities, and switch counts. Decentralized actors are trained with a centralized critic for stable on-policy learning, while a physics-informed OpenDSS environment enforces electrical feasibility. Experiments on IEEE 123-bus and 8500-node feeders show HAPPO outperforms PPO, QMIX, Mean-Field RL, and other baselines in restored power, convergence stability, and multi-seed reproducibility. Under a 2400 kW generation cap, the framework restores over 95\% of available load on both systems with low-latency execution, supporting practical real-time PDS restoration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。