arXiv:2605.12462cs.AIcs.CY2026-05被引 1

构建电力需求响应仿真环境,助力电网优化用户用电决策

Towards Affordable Energy: A Gymnasium Environment for Electric Utility Demand-Response Programs

论文配图:Towards Affordable Energy: A Gymnasium Environment for Electric Utility Demand-Response Programs
图 1 · 摘自论文原文
  • 设计可兼容Gymnasium的市场级仿真环境,模拟电网与用户互动
  • 集成真实极端事件电价模型与建筑能耗物理特征
  • 支持多目标奖励函数,适用于不同电网运营策略训练

极端天气和波动的批发市场使居民面临重大财务风险,但配电层级的需求响应仍为被充分使用的电网灵活性工具。尽管在高电价时段发放金融激励可保护消费者,但如何优化这一序列决策过程对强化学习构成独特挑战,即便存在大量公开的离线智能电表和批发电价数据。现有离线数据无法捕捉电力公司定价信号与用户接受度及适应性之间的动态反馈循环。为此,我们提出DR-Gym,一个开源、在线的Gymnasium兼容环境,用于从电力公司的视角训练和评估需求响应。不同于现有的设备级能源仿真器,本环境聚焦于市场级电力公司场景,提供丰富的观测空间。仿真器包含基于真实极端事件校准的分段式电价模型,以及基于物理的建筑负荷曲线。学习信号采用可配置的多目标奖励函数,以支持多样化学习目标。通过基准策略与数据快照,我们验证了该仿真器生成现实且可学习环境的能力。

原文摘要 · Abstract (English)

Extreme weather and volatile wholesale electricity markets expose residential consumers to catastrophic financial risks, yet demand response at the distribution level remains an underutilized tool for grid flexibility and energy affordability. While a demand-response program can shield consumers by issuing financial credits during high-price periods, optimizing this sequential decision-making process presents a unique challenge for reinforcement learning despite the plentiful offline historical smart meter and wholesale pricing data available publicly. Offline historical data fails to capture the dynamic, interactive feedback loop between an electric utility's pricing signals and customer acceptance and adaptation to a demand-response program. To address this, we introduce DR-Gym, an open-source, online Gymnasium-compatible environment designed to train and evaluate demand-response from the electric utility's perspective. Unlike existing device-level energy simulators, our environment focuses on the market-level electric utility setting and provides a rich observational space relevant to the electric utility. The simulator additionally features a regime-switching wholesale price model calibrated to real-world extreme events, alongside physics-based building demand profiles. For our learning signal, we use a configurable, multi-objective reward function for specifying diverse learning objectives. We demonstrate through baseline strategies and data snapshots the capability of our simulator to create realistic and learnable environments.

需求响应强化学习电力系统仿真环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。