arXiv:2509.23620eess.SYcs.LG2025-09中稿 · publication in IEE…被引 1

针对电力系统广域阻尼控制中的通信延迟问题,提出风险约束强化学习新方法。

Communication-aware Wide-Area Damping Control using Risk-Constrained Reinforcement Learning

  • 引入均值-方差风险约束,改进传统LQR以应对通信不确定性
  • 在IEEE 68节点系统上验证,相比传统补偿方法性能更优
  • 适合关注电网安全与通信鲁棒性的控制工程师

非理想通信链路(尤其是延迟)严重影响电力系统广域阻尼控制(WADC)的快速响应。传统方法依赖高精度延迟估计与补偿,但难以应对链路故障或网络扰动等其他网络问题。本文提出一种风险约束框架,可有效处理通信延迟及其他不确定性。所提WADC模型包含同步发电机(SGs)和电压源变流器(VSCs),通过引入均值-方差风险约束改进线性二次型调节器(LQR)成本函数。采用基于随机梯度下降与最大值查询器(SGDmax)的强化学习算法求解,证明其在零阶策略梯度下仍能以高概率收敛至稳定点。在IEEE 68节点系统上的数值实验验证了算法收敛性、VSCs的阻尼能力,并表明该方法在延迟估计误差存在时优于传统补偿方法。本设计不仅提升大延迟下的性能,还能有效抑制最坏情况振荡,适用于多种通信问题与网络扰动。

原文摘要 · Abstract (English)

Non-ideal communication links, especially delays, critically affect fast networked controls in power systems, such as the wide-area damping control (WADC). Traditionally, a delay estimation and compensation approach is adopted to address this cyber-physical coupling, but it demands very high accuracy for the fast WADC and cannot handle other cyber concerns like link failures or {cyber perturbations}. Hence, we propose a new risk-constrained framework that can target the communication delays, yet amenable to general uncertainty under the cyber-physical couplings. Our WADC model includes the synchronous generators (SGs), and also voltage source converters (VSCs) for additional damping capabilities. To mitigate uncertainty, a mean-variance risk constraint is introduced to the classical optimal control cost of the linear quadratic regulator (LQR). Unlike estimating delays, our approach can effectively mitigate large communication delays by improving the worst-case performance. A reinforcement learning (RL)-based algorithm, namely, stochastic gradient-descent with max-oracle (SGDmax), is developed to solve the risk-constrained problem. We further show its guaranteed convergence to stationarity at a high probability, even using the simple zero-order policy gradient (ZOPG). Numerical tests on the IEEE 68-bus system not only verify SGDmax's convergence and VSCs' damping capabilities, but also demonstrate that our approach outperforms conventional delay compensator-based methods under estimation error. While focusing on performance improvement under large delays, our proposed risk-constrained design can effectively mitigate the worst-case oscillations, making it equally effective for addressing other communication issues and cyber perturbations.

电力系统强化学习阻尼控制风险约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。