arXiv:2505.10296cs.LG2025-05被引 11

用分层强化学习优化电动公交充电,应对行程与电价不确定性。

Optimizing Electric Bus Charging Scheduling with Uncertainties Using Hierarchical Deep Reinforcement Learning

  • 分两层决策:高层选充电时机,底层定充电功率,实现多时间尺度调度。
  • 在真实数据集上,充电成本降低12.3%,且支持千辆级车队调度。
  • 适合城市交通管理者和新能源公交运营商参考使用。

电动公交车(EBs)的普及是可持续发展的重要一步。借助物联网(IoT)系统,充电桩可基于实时数据自主制定充电计划。然而,由于行程时间、能耗及电价波动带来的不确定性,优化电动公交车充电调度仍具挑战性。此外,实际应用中需在多时间尺度上高效决策,并具备大规模车队的可扩展性。本文提出一种分层深度强化学习(HDRL)方法,将原始马尔可夫决策过程(MDP)重构为两个增强型MDP。为此,设计新型算法双演员-评论家多智能体近端策略优化增强版(DAC-MAPPO-E)。针对双演员-评论家(DAC)算法在大规模车队中的可扩展性问题,在高层重新设计去中心化演员网络并引入注意力机制,提取各电动公交相关的全局状态信息,缩小神经网络规模;在低层将多智能体近端策略优化(MAPPO)集成至DAC框架,实现去中心化且协同的充电功率决策,降低计算复杂度并提升收敛速度。基于真实世界数据的大量实验表明,该方法在优化电动公交车队充电调度方面具有显著性能优势与良好可扩展性。

原文摘要 · Abstract (English)

The growing adoption of Electric Buses (EBs) represents a significant step toward sustainable development. By utilizing Internet of Things (IoT) systems, charging stations can autonomously determine charging schedules based on real-time data. However, optimizing EB charging schedules remains a critical challenge due to uncertainties in travel time, energy consumption, and fluctuating electricity prices. Moreover, to address real-world complexities, charging policies must make decisions efficiently across multiple time scales and remain scalable for large EB fleets. In this paper, we propose a Hierarchical Deep Reinforcement Learning (HDRL) approach that reformulates the original Markov Decision Process (MDP) into two augmented MDPs. To solve these MDPs and enable multi-timescale decision-making, we introduce a novel HDRL algorithm, namely Double Actor-Critic Multi-Agent Proximal Policy Optimization Enhancement (DAC-MAPPO-E). Scalability challenges of the Double Actor-Critic (DAC) algorithm for large-scale EB fleets are addressed through enhancements at both decision levels. At the high level, we redesign the decentralized actor network and integrate an attention mechanism to extract relevant global state information for each EB, decreasing the size of neural networks. At the low level, the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm is incorporated into the DAC framework, enabling decentralized and coordinated charging power decisions, reducing computational complexity and enhancing convergence speed. Extensive experiments with real-world data demonstrate the superior performance and scalability of DAC-MAPPO-E in optimizing EB fleet charging schedules.

电动公交强化学习调度优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。