用图学习方法优化电力与交通网络协同抢修,提升恢复效率。
Learning-aided Bigraph Matching Approach to Multi-Crew Restoration of Damaged Power Networks Coupled with Road Transportation Networks
- 构建电力与交通双图融合模型,通过图强化学习设计任务分配激励函数。
- 在8500节点电网上测试,性能比随机策略高3倍,恢复速度更快。
- 适合电力系统韧性评估、应急调度研究者参考。
自然灾害等事件后,关键基础设施网络的韧性取决于恢复速度和功能复原程度。资源分配是组合优化问题,需决定由哪些抢修队伍修复哪些节点及顺序。本文提出一种新型图模型,将人员与交通节点和电网节点整合为异构图。结合图强化学习(GRL)与二部图匹配,利用图神经网络(基于近端策略优化和神经演化训练)设计激励函数,实现对环境状态的抽象建模,确保跨故障场景的泛化能力。该激励函数用于加权二部图匹配,完成人员到任务的分配。使用预计算最优路径的高效仿真环境进行训练。以包含8500节点配电网络和21平方公里交通网络的IEEE测试案例为基础,覆盖不同受损节点数、调度点和人员数量的场景。结果表明,所提方法在多种场景下具备良好泛化性与可扩展性,学习策略性能较随机策略提升3倍,且在计算时间(多个数量级)和恢复电量方面均优于传统优化方法。
原文摘要 · Abstract (English)
The resilience of critical infrastructure networks (CINs) after disruptions, such as those caused by natural hazards, depends on both the speed of restoration and the extent to which operational functionality can be regained. Allocating resources for restoration is a combinatorial optimal planning problem that involves determining which crews will repair specific network nodes and in what order. This paper presents a novel graph-based formulation that merges two interconnected graphs, representing crew and transportation nodes and power grid nodes, into a single heterogeneous graph. To enable efficient planning, graph reinforcement learning (GRL) is integrated with bigraph matching. GRL is utilized to design the incentive function for assigning crews to repair tasks based on the graph-abstracted state of the environment, ensuring generalization across damage scenarios. Two learning techniques are employed: a graph neural network trained using Proximal Policy Optimization and another trained via Neuroevolution. The learned incentive functions inform a bipartite graph that links crews to repair tasks, enabling weighted maximum matching for crew-to-task allocations. An efficient simulation environment that pre-computes optimal node-to-node path plans is used to train the proposed restoration planning methods. An IEEE 8500-bus power distribution test network coupled with a 21 square km transportation network is used as the case study, with scenarios varying in terms of numbers of damaged nodes, depots, and crews. Results demonstrate the approach's generalizability and scalability across scenarios, with learned policies providing 3-fold better performance than random policies, while also outperforming optimization-based solutions in both computation time (by several orders of magnitude) and power restored.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。