用强化学习优化电动车充电,降本增效还可靠。
Reinforcement Learning-based Approach for Vehicle-to-Building Charging with Heterogeneous Agents and Long Term Rewards
- 结合DDPG与动作掩码,处理连续动作空间和多智能体决策
- 在真实数据上实现显著成本降低,且满足全部充电需求
- 适合研究智能电网、电动车管理的从业者参考
将电动汽车电池作为储能资源进行策略聚合,可优化电网用电需求,尤其适用于提供工作场所充电的大型办公楼。该任务需在长时间周期(如一个月)内优化充放电行为,以降低峰值电费和净峰值负荷,涉及在不确定性下做出序列决策,面临奖励延迟且稀疏、连续动作空间以及跨多种条件泛化能力不足等问题。现有方法如启发式策略难以应对动态环境中的实时决策,传统强化学习模型则受限于大规模状态-动作空间、多智能体场景及长期奖励优化需求。为此,本文提出一种新强化学习框架,融合深度确定性策略梯度(DDPG)、动作掩码机制与高效混合整数线性规划(MILP)驱动的策略引导。该方法在探索连续动作空间时兼顾用户充电需求。基于某主要电动车制造商的真实数据验证表明,本方法全面优于多个成熟基线与可扩展启发式方法,在满足所有充电要求的同时实现显著成本节约。结果表明,该方法是首个可扩展且具泛化的电动车辆到建筑(V2B)能源管理解决方案。
原文摘要 · Abstract (English)
Strategic aggregation of electric vehicle batteries as energy reservoirs can optimize power grid demand, benefiting smart and connected communities, especially large office buildings that offer workplace charging. This involves optimizing charging and discharging to reduce peak energy costs and net peak demand, monitored over extended periods (e.g., a month), which involves making sequential decisions under uncertainty and delayed and sparse rewards, a continuous action space, and the complexity of ensuring generalization across diverse conditions. Existing algorithmic approaches, e.g., heuristic-based strategies, fall short in addressing real-time decision-making under dynamic conditions, and traditional reinforcement learning (RL) models struggle with large state-action spaces, multi-agent settings, and the need for long-term reward optimization. To address these challenges, we introduce a novel RL framework that combines the Deep Deterministic Policy Gradient approach (DDPG) with action masking and efficient MILP-driven policy guidance. Our approach balances the exploration of continuous action spaces to meet user charging demands. Using real-world data from a major electric vehicle manufacturer, we show that our approach comprehensively outperforms many well-established baselines and several scalable heuristic approaches, achieving significant cost savings while meeting all charging requirements. Our results show that the proposed approach is one of the first scalable and general approaches to solving the V2B energy management challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。