arXiv:2605.04555cs.LGcs.SY2026-05被引 2

用反事实模型让空调控制强化学习只需5周数据,比之前快3倍。

Counter-Dyna: Data-Efficient RL-Based HVAC Control using Counterfactual Building Models

  • 构建基于状态空间不变性的反事实仿真模型,只关注受调控影响的部分。
  • 在模拟中仅用5周数据达到17%的节能效果,显著少于之前的6-12个月。
  • 适合想快速落地强化学习的楼宇能源管理研究者和工程师。

基于模型的强化学习(MBRL)为建筑节能管理提供了高效的数据利用方案,结合预测建模与强化学习优势。尽管已有方法降低了训练所需数据量,但仍需数月环境交互才能获得满意控制策略。关键原因在于现有代理模型试图预测全状态空间,包括不受控制影响的天气和电价,或完全忽略这些变量。为此,我们提出Counter-Dyna,通过利用状态空间中的不变性,提升Dyna这一MBRL方法的数据效率。我们构建了数据高效的反事实代理模型(CSM),并将其用于Dyna,显著加快了强化学习训练速度。相比以往需6-12个月环境交互的最先进方法,本方法仅需5周。我们在标准BOPTEST框架下,使用近端策略优化算法(PPO)进行大规模仿真评估,结果显示在假设部署场景中可实现5.3%至17.0%的成本节约。该工作是推动强化学习在暖通空调控制中实际应用的重要一步。

原文摘要 · Abstract (English)

Model-based reinforcement learning (MBRL) offers a promising approach for data-efficient energy management in buildings, combining the strengths of predictive modeling and reinforcement learning. While previous MBRL methods applied to HVAC control have reduced training data requirements, they still require several months of interaction with the building to learn a satisfactory control policy. A key reason is that existing surrogate models attempt to predict the entire state-space, including weather and electricity prices that are unaffected by control actions, or completely ignore these variables. Addressing these issues, we propose Counter-Dyna, a method that enhances the data-efficiency of Dyna, an MBRL method. We create data-efficient counterfactual surrogate models (CSM) by leveraging invariances in the state-space. Using a CSM in Dyna speeds up RL training measured in environment interaction data compared to previous results. In comparison with previous state-of-the-art that used 6-12 months of environment interactions, our method needs only 5 weeks. We evaluate our method in a large simulation study using the literature standard BOPTEST framework and proximal policy algorithm (PPO) as the RL algorithm. Our results show cost-saving potentials of 5.3% to 17.0% in a hypothetical deployment scenario. Our work is a significant step towards making real-world deployment of RL algorithms in HVAC control practically viable.

强化学习空调控制节能优化数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。