arXiv:2608.00679cs.AIcs.MA2026-08

用物理约束指导安全干预,实现大规模电动车充电的高效协调。

HetGPS: Scalable Graph Multi-Agent Reinforcement Learning with Physics-Anchored Adaptive Safety for EV Charging

论文配图:HetGPS: Scalable Graph Multi-Agent Reinforcement Learning with Physics-Anchored Adaptive Safety for EV Charging
图 1 · 摘自论文原文
  • 分离干预强度与方向,用图模型学风险,用物理模型定修正方向。
  • 在200至3218辆电动车场景下,电压越限率从7.74%降至3.44%,出发成功率超99%。
  • 模型参数规模不随车队扩大而增长,可零样本迁移至不同系统。

针对大规模网络耦合智能体的安全干预问题,需在保护共享约束的同时避免过度干扰任务决策。本文提出HetGPS,一种融合学习图风险与物理锚定修正的混合图控框架,将干预幅度与修正方向分离。动作条件图残差模型调度状态依赖的干预权限,而物理模型决定修正方向。在电动车充电场景中,该框架与参数共享的异构图软演员-评论家策略结合,实现拓扑感知协调,且模型规模独立于车队大小。在包含200至3,218辆电动车的五个嵌套配电网络上,自适应权限将母线-步电压越限率从3.93%–7.74%降至0.52%–3.44%,同时保持99.06%–100%的离站成功率。相较固定权限的物理投影方法,其在所有五组网络上提升平均奖励,在四组降低平均安全评分。部署的策略与风险模型在各尺度均含383,702个可学习参数;在3,218辆电动车时,对应集中式SAC Actor约大170倍。在八变压器系统训练的策略可零样本迁移至十六和三十二变压器系统,电压越限率0.57%–0.75%,离站成功率不低于99.99%。结果表明,学习图风险可在大规模下分配干预权限,而馈线物理特性锚定修正动作。

原文摘要 · Abstract (English)

Safety interventions for large populations of network-coupled agents must protect shared constraints without unnecessarily overriding task-oriented policy decisions. We present HetGPS, a hybrid graph-control framework synergizing learned graph risk with physics-anchored correction by separating intervention magnitude from corrective direction. An action-conditioned graph residual model schedules state-dependent intervention authority, while a physics model determines its direction. For electric vehicle (EV) charging, we couple this filter with a parameter-shared heterogeneous graph soft actor-critic policy, enabling topology-aware coordination with a learned model size independent of fleet size. Across five nested distribution networks with 200--3,218 EVs and 100 evaluation days, Adaptive Authority reduces bus--step voltage violations from 3.93--7.74\% without filtering to 0.52--3.44\%, while maintaining 99.06--100\% departure success. Relative to the same physics-directed projection with fixed authority, it improves mean reward on all five networks and lowers the mean safety score on four. The deployed policy-and-risk model contains 383,702 learned parameters at every scale; at 3,218 EVs, a matched centralized SAC actor is about $170\times$ larger. A policy trained on the eight-transformer system transfers zero-shot to the 16- and 32-transformer systems, attaining 0.57--0.75\% violation rates and at least 99.99\% departure success. These results show that learned graph risk can allocate intervention authority at scale while feeder physics anchors corrective action.

多智能体强化学习电动车充电安全控制图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。