arXiv:2409.19716cs.LGcs.AI2024-09中稿 · European Control C…被引 2

用约束强化学习优化暖气机控制,兼顾节能与舒适度。

Constrained Reinforcement Learning for Safe Heat Pump Control

  • 提出新建筑仿真器I4B和基于光滑对数屏障的约束算法CSAC-LB
  • 在少数据下实现高能效与95%以上舒适度满足率
  • 适合智能楼宇控制、能源优化领域研究人员参考

约束强化学习(Constrained Reinforcement Learning, RL)已成为强化学习的重要研究方向,将约束与奖励结合对提升各类控制任务的安全性与性能至关重要。在建筑供暖系统中,优化能耗同时维持居民热舒适度可自然建模为约束优化问题。然而,采用RL求解需大量数据,因此需要高精度且通用的仿真器。本文提出新型建筑仿真器I4B,支持多种使用场景,并应用无模型约束强化学习算法——带线性平滑对数屏障函数的约束软演员-评论家算法(CSAC-LB)解决供暖优化问题。与基线算法对比表明,CSAC-LB在数据探索效率、约束满足率和性能方面表现优异。实验验证最优运行通常位于舒适度边界附近,且该算法对传感器噪声和模型失配具有鲁棒性。

原文摘要 · Abstract (English)

Constrained Reinforcement Learning (RL) has emerged as a significant research area within RL, where integrating constraints with rewards is crucial for enhancing safety and performance across diverse control tasks. In the context of heating systems in the buildings, optimizing the energy efficiency while maintaining the residents' thermal comfort can be intuitively formulated as a constrained optimization problem. However, to solve it with RL may require large amount of data. Therefore, an accurate and versatile simulator is favored. In this paper, we propose a novel building simulator I4B which provides interfaces for different usages and apply a model-free constrained RL algorithm named constrained Soft Actor-Critic with Linear Smoothed Log Barrier function (CSAC-LB) to the heating optimization problem. Benchmarking against baseline algorithms demonstrates CSAC-LB's efficiency in data exploration, constraint satisfaction and performance.

强化学习供暖控制约束学习建筑仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。