arXiv:2605.11102cs.LGcs.AI2026-05

用强化学习优化电力系统初始解,显著提升复杂工况下的收敛性。

Newton's Lantern: A Reinforcement Learning Framework for Finetuning AC Power Flow Warm Start Models

论文配图:Newton's Lantern: A Reinforcement Learning Framework for Finetuning AC Power Flow Warm Start Models
图 1 · 摘自论文原文
  • 基于扰动预测构建奖励模型,以迭代次数为监督信号进行微调
  • 在多个电网规模测试中实现100%收敛且平均迭代次数最少
  • 解决传统方法在电压崩溃附近失效的问题,适合电力系统优化场景

神经网络预热解可大幅减少求解交流潮流问题所需的牛顿-拉夫逊迭代次数,但现有监督方法在重载工况、接近电压崩溃时泛化能力差。我们证明了牛顿-拉夫逊迭代次数存在下界,该下界取决于初始误差的方向而非大小,且当潮流雅可比矩阵最小奇异值趋近零时,该下界变得无效,揭示了监督回归在鞍结分岔附近的失效机制。受此分析启发,我们提出Newton's Lantern微调框架,结合组相对策略优化与基于基础模型预测扰动的自学习奖励模型,直接以迭代次数作为监督信号。在IEEE 118节点、GOC 500节点和GOC 2000节点基准测试中,Newton's Lantern是唯一能在所有测试快照上收敛并达到最小均值迭代次数的方法。

原文摘要 · Abstract (English)

Neural warm starts can sharply reduce the number of Newton-Raphson iterations required to solve the AC power flow problem, but existing supervised approaches generalize poorly on heavily loaded instances near voltage collapse. We prove a lower bound on the Newton-Raphson iteration count that depends on the direction of the warm start error rather than on its magnitude, and show as a corollary that the bound becomes vacuous as the smallest singular value of the power-flow Jacobian shrinks, identifying the failure mode of supervised regression near the saddle-node bifurcation. Motivated by this analysis, we introduce Newton's Lantern, a finetuning pipeline that combines group relative policy optimization with a learned reward model trained on perturbations of the base model's predictions, using the iteration count itself as the supervisory signal. Across IEEE 118-bus, GOC 500-bus, and GOC 2000-bus benchmarks, Newton's Lantern is the only method that converges on every test snapshot while attaining the smallest mean iteration count.

电力系统强化学习潮流计算微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。