arXiv:2604.24549cs.LGcs.AI2026-04

提出无需通信的分布式学习方法,高效优化电网边缘设备运行。

GradMAP: Gradient-Based Multi-Agent Proximal Learning for Grid-Edge Flexibility

论文配图:GradMAP: Gradient-Based Multi-Agent Proximal Learning for Grid-Edge Flexibility
图 1 · 摘自论文原文
  • 每台设备独立训练神经网络策略,仅用本地观测决策
  • 15分钟内完成1000台设备训练,较基准快3-5倍
  • 适合大规模电网边缘设备协同优化场景

协调大量电网边缘设备需在部署时保持完全去中心化,同时遵守三相交流配电网物理规律。本文提出基于梯度的多智能体近端学习(GradMAP),为每个智能体独立训练神经网络策略,不共享参数,且在线决策时仅依赖本地观测,无需通信。离线训练中,GradMAP将可微分三相交流潮流模型嵌入原始-对偶学习循环,并利用隐式微分将精确的网络约束违规信息反向传播以更新策略参数。为加速训练,其在更直接的策略输出(动作)空间而非概率分布空间定义信任区域,通过近端代理复用昂贵的环境梯度。在包含1000个管理电池、热泵和可控发电机的智能体的IEEE 123节点馈线案例研究中,GradMAP在单块工作站级NVIDIA RTX PRO 5000 Blackwell 48GB GPU上15分钟内学习到去中心化策略,显著降低三相交流潮流约束违规。相比基于梯度的自监督学习基准,训练速度提升3–5倍,远超多智能体强化学习基准。样本外测试中,该方法亦实现最低运行成本与约束违规。

原文摘要 · Abstract (English)

Coordinating large populations of grid-edge devices requires learning methods that remain fully decentralised in deployment while still respecting three-phase AC distribution-network physics. This paper proposes gradient-based multi-agent proximal learning (GradMAP) to address this challenge. GradMAP trains independent neural-network policies for each agent without any parameter sharing, and each agent uses only its own local observation for online decision-making without communication. During offline training, GradMAP embeds a differentiable three-phase AC power-flow model in a primal-dual learning loop and uses implicit differentiation to propagate exact network-constraint violations to update the policy parameters. To speed up training, GradMAP reuses expensive environment gradients through a proximal surrogate within a trust region defined in the more direct policy-output (action) space, instead of the probability distribution space used in other works, such as PPO. In case studies with 1,000 agents managing batteries, heat pumps, and controllable generators on the IEEE 123-bus feeder, GradMAP learns decentralised policies that minimise three-phase AC load-flow constraint violations within 15 minutes of training on a single workstation-class NVIDIA RTX PRO 5000 Blackwell 48GB GPU. This is a 3--5x training speed-up over gradient-based self-supervised learning benchmarks and substantially better training efficiency than multi-agent reinforcement-learning benchmarks. In out-of-sample tests, GradMAP also delivers among the lowest operating cost and constraint violations.

电网优化多智能体分布式学习三相潮流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。