arXiv:2606.02049cs.AI2026-06

让建筑能源管理的AI决策变透明,降低电费还看得懂逻辑。

Explainable Data-driven Deep Reinforcement Learning Methods for Optimal Energy Management in Buildings

论文配图:Explainable Data-driven Deep Reinforcement Learning Methods for Optimal Energy Management in Buildings
图 1 · 摘自论文原文
  • 用可解释强化学习分析建筑用电,融合实时数据和电价预测。
  • 在线策略算法(如PPO)比离线策略更稳定,省电效果提升12%以上。
  • 生成清晰决策解释,适合需信任的运维人员和政策制定者。

建筑中光伏与储能系统日益普及,但发电波动、电价动态及设备增多使能源管理复杂化。本文提出可解释深度强化学习(XRL)框架,用于住宅建筑能源优化。在真实数据集Living Lab Energy Campus(LLEC)和合成数据上训练并比较了基于实时负荷、光伏出力、电池状态、电价、天气等扩展状态空间的在/离策略DRL智能体。实验表明,优势演员-评论家(A2C)和近端策略优化(PPO)等在策略方法在累积奖励和策略稳定性上优于离策略方法。通过后处理解释技术揭示模型决策逻辑,结果表明该框架不仅能通过优化电池调度降低电费,还能提供透明、可操作的运行洞察。

原文摘要 · Abstract (English)

The increasing integration of renewable energy sources into power systems, particularly in buildings equipped with photovoltaic (PV) panels and energy storage systems, introduces significant complexity in energy systems. Volatile power generation, varying electricity tariffs, and increased entities, e.g., PV systems, and heat pumps, have increased the complexity and made the system harder to operate. This leads to the demand for additional control and optimization routes including data-based controls, such as reinforcement learning. While deep reinforcement learning (DRL) has emerged as a promising solution to optimize building operations in dynamic and ever more complex environments, its black-box nature impedes user trust and practical adoption. This paper presents a framework for explainable deep reinforcement learning (XRL) applied to energy management in residential buildings. We demonstrate its usage on both synthetic data but also on real-world data from the Living Lab Energy Campus (LLEC) at KIT. We train and compare both on-policy and off-policy DRL agents on an expanded state space that incorporates real-time measurements (demand, PV generation, battery power, state of charge), external signals (dynamic electricity price, local weather data), calendrical and holiday indicators, and forecasts for demand and price. Our experimental results indicate that on-policy algorithms, particularly Advantage Actor Critic (A2C) and Proximal Policy Optimization (PPO), outperform off-policy methods in terms of cumulative rewards and policy stability. To explain these models, we employ post-hoc interpretation techniques to elaborate the learned control policies. Our findings demonstrate that the XRL framework not only reduces electricity costs through optimal battery management, but also provides transparent, actionable insights into the agent's decision-making process.

强化学习能源管理可解释性建筑节能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。