用强化学习协调充电桩,在信息有限下既保电压安全又降成本。
Safe Decentralized Operation of EV Virtual Power Plant with Limited Network Visibility via Multi-Agent Reinforcement Learning

- 多智能体强化学习,中心化训练、分布式执行,约束电压与充电需求。
- 33节点电网实验显示电压越限减少45%,运营成本降低10%。
- 适合电力系统、智能交通领域,关注新能源消纳与配网安全的人看。
随着电网向净零目标迈进,用户侧可再生能源推动分布式能源资源(DERs)快速增长。虚拟电厂(VPP)日益用于协调这些资源以支持配电网络(PDN)运行,其中电动汽车充电站(EVCS)因其对局部电压的显著影响成为关键资产。然而,实践中VPP只能基于配电系统运营商提供的有限聚合信息做出运营决策,缺乏完整的网络状态可视性。本文提出一种安全增强型VPP框架,用于在真实信息受限条件下协调多个EVCS,确保电压安全的同时维持经济运行。我们开发了基于Transformer的拉格朗日多智能体近端策略优化(TL-MAPPO),其中各EVCS智能体通过中心化训练并采用拉格朗日正则化学习去中心化充电策略,以强制满足电压和负荷需求约束。每个EVCS智能体部署基于Transformer的嵌入层,捕捉电价、负荷与充电需求间的时序相关性,提升决策质量。在真实的33节点配电网络上实验表明,相比代表性多智能体深度强化学习基线,该框架使电压越限减少约45%,运营成本降低约10%,凸显其在实际VPP部署中的潜力。
原文摘要 · Abstract (English)
As power systems advance toward net-zero targets, behind-the-meter renewables are driving rapid growth in distributed energy resources (DERs). Virtual power plants (VPPs) increasingly coordinate these resources to support power distribution network (PDN) operation, with EV charging stations (EVCSs) emerging as a key asset due to their strong impact on local voltages. However, in practice, VPPs must make operational decisions with only partial visibility of PDN states, relying on limited, aggregated information shared by the distribution system operator. This work proposes a safety-enhanced VPP framework for coordinating multiple EVCSs under such realistic information constraints to ensure voltage security while maintaining economic operation. We develop Transformer-assisted Lagrangian Multi-Agent Proximal Policy Optimization (TL-MAPPO), in which EVCS agents learn decentralized charging policies via centralized training with Lagrangian regularization to enforce voltage and demand-satisfaction constraints. A transformer-based embedding layer deployed on each EVCS agent captures temporal correlations among prices, loads, and charging demand to improve decision quality. Experiments on a realistic 33-bus PDN show that the proposed framework reduces voltage violations by approximately 45% and operational costs by approximately 10% compared to representative multi-agent DRL baselines, highlighting its potential for practical VPP deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。