arXiv:2602.19223cs.AIcs.LG2026-02被引 1

为城市能源控制设计多指标基准,揭示MARL算法真实表现

Characterizing MARL for Energy Control: A Multi-KPI Benchmark on the CityLearn Environment

  • 在CityLearn环境中对比多种MARL算法与训练策略
  • DTDE策略在平均与最差情况下均优于CTDE,且电池寿命更可持续
  • 提出新指标评估建筑贡献与储能寿命,适合关注实际部署的读者

城市能源系统的优化对建设可持续、韧性的智慧城市至关重要,而系统复杂性带来多决策单元的协调挑战。本文以CityLearn为案例环境,构建首个涵盖多个关键性能指标(KPI)的全面基准,用于评估多智能体强化学习(MARL)在能源管理中的表现。该环境真实模拟城市能源系统,集成多种储能设备与可再生能源。研究采用主流基线如PPO和SAC,对比去中心化训练-去中心化执行(DTDE)与集中式训练-去中心化执行(CTDE)等不同训练方案及神经网络架构。实验发现,DTDE在平均与最差情形下均表现更优;时间依赖学习显著提升对爬坡率与电池使用等时序相关指标的控制能力,促进电池可持续运行。此外,政策对智能体或资源移除表现出强鲁棒性,验证其抗干扰能力与去中心化优势。研究还引入新指标,解决实际部署中建筑贡献度与电池寿命评估难题。

原文摘要 · Abstract (English)

The optimization of urban energy systems is crucial for the advancement of sustainable and resilient smart cities, which are becoming increasingly complex with multiple decision-making units. To address scalability and coordination concerns, Multi-Agent Reinforcement Learning (MARL) is a promising solution. This paper addresses the imperative need for comprehensive and reliable benchmarking of MARL algorithms on energy management tasks. CityLearn is used as a case study environment because it realistically simulates urban energy systems, incorporates multiple storage systems, and utilizes renewable energy sources. By doing so, our work sets a new standard for evaluation, conducting a comparative study across multiple key performance indicators (KPIs). This approach illuminates the key strengths and weaknesses of various algorithms, moving beyond traditional KPI averaging which often masks critical insights. Our experiments utilize widely accepted baselines such as Proximal Policy Optimization (PPO) and Soft Actor Critic (SAC), and encompass diverse training schemes including Decentralized Training with Decentralized Execution (DTDE) and Centralized Training with Decentralized Execution (CTDE) approaches and different neural network architectures. Our work also proposes novel KPIs that tackle real world implementation challenges such as individual building contribution and battery storage lifetime. Our findings show that DTDE consistently outperforms CTDE in both average and worst-case performance. Additionally, temporal dependency learning improved control on memory dependent KPIs such as ramping and battery usage, contributing to more sustainable battery operation. Results also reveal robustness to agent or resource removal, highlighting both the resilience and decentralizability of the learned policies.

MARL能源控制多指标评估城市智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。