arXiv:2607.05553cs.LGcs.SY2026-07中稿 · IEEE SmartGridComm…被引 1

用联邦强化学习提升电网故障后稳定控制,100%成功且响应快72.4%。

Federated Physics-Grounded Reinforcement Learning for Distributed Stability Control in Smart Grids

论文配图:Federated Physics-Grounded Reinforcement Learning for Distributed Stability Control in Smart Grids
图 1 · 摘自论文原文
  • 将电网稳定控制转为基于物理耦合关系的多智能体强化学习。
  • 在IEEE 39节点系统上实现24次试验全部稳定,平均恢复时间缩短72.4%。
  • 无需中心协调,每节点推理延迟满足实时通信标准,适合分布式部署。

智能电网的暂态稳定控制需快速抑制故障后的发电机频率与转子角偏差,防止级联失效。本文提出FedPPO-PG——一种基于物理约束邻域的联邦多智能体近端策略优化框架,将暂态稳定控制重构为直接优化闭环稳定性目标的协作式多智能体强化学习问题。每个发电机组部署独立本地策略,其状态包含从故障后克罗恩简化电抗矩阵中识别出的两个最强耦合电气邻居的频率偏差。通过引导式策略初始化,所有智能体从经典分散控制器出发预热;在集中训练-分散执行(CTDE)范式下,由中央评价者指导优势估计。在IEEE 39节点基准系统上,对五组训练及三组未见故障工况进行仿真评估,FedPPO-PG在全部24次试验中实现100%稳定,平均稳定时间减少72.4%,控制功率降低7–14倍,优于集中式基线方法。部署时各智能体独立执行,无需中央协调器,单智能体推理延迟满足IEEE/IEC 60255-118-1-2018实时报告要求。

原文摘要 · Abstract (English)

Transient stability control in smart grids requires rapid post-fault damping of generator frequency and rotor angle deviations to prevent cascading failures. This paper proposes FedPPO-PG, a Federated Multi-Agent Proximal Policy Optimization framework with Physics-Grounded neighborhoods, which reformulates transient stability control as a cooperative multi-agent reinforcement learning problem optimized directly against closed-loop stability objectives. Each generator hosts an independent local actor augmented with the frequency deviations of its two most strongly coupled electrical neighbors, identified from the post-fault Kron-reduced susceptance matrix. A guided policy initialization phase warm-starts all actors from the classical decentralized controller, while a centralized critic guides advantage estimation under the centralized training--decentralized execution (CTDE) paradigm. Evaluated on a simulation of the IEEE 39-bus benchmark system across five training and three unseen fault contingencies, FedPPO-PG achieves 100% stabilization in all 24 trials, reduces mean stability time by 72.4%, and cuts the control power by 7-14 times compared to the centralized baseline. Each actor executes independently with no central coordinator at deployment, and the per-actor inference latency satisfies the IEEE/IEC 60255-118-1-2018 real-time reporting requirements.

电网控制强化学习联邦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。