对比强化学习与模型预测控制在住宅空调中的实际部署效果。
Comparative Field Deployment of Reinforcement Learning and Model Predictive Control for Residential HVAC
- 用强化学习和模型预测控制分别优化空调温控策略。
- 两者节能均超18%,但强化学习初期导致住户感觉偏冷。
- 强化学习部署更省工程量,但需解决初始化与状态偏差问题。
模型预测控制(MPC)在住宅供暖、通风与空调(HVAC)系统中已证明显著优于现有控制方法,但部署常需大量工程投入。强化学习(RL)可能实现相近性能且部署更简便,但其在住宅HVAC中的实际应用仍缺乏验证,尤其在用户舒适度和数据需求方面存在疑问。为此,我们在寒冷气候的一户人家中,各部署一个MPC变体和一个基于模型的RL变体,持续一个月。控制器根据室内温度和制热用电量调整空气源热泵的设定温度。相比固定设定值运行,MPC节省了18.1%(95%置信区间:4.4至30.9%)的天气归一化热泵能耗,而RL节省20.9%(2.6至38.3%)。MPC维持了可接受的用户舒适度,而RL使房屋整体较冷,尤其在初始适应阶段,导致三次用户不适反馈。两种算法的数据需求相似。我们估算,在另一栋房屋中进行全新部署时,RL所需工程工作量约为MPC的三分之一。尽管RL降低了部署负担,但仍面临安全控制器初始化及模型与真实状态/动作空间不匹配等挑战。
原文摘要 · Abstract (English)
Model Predictive Control (MPC) has demonstrated significant performance improvements over today's control methods for residential Heating, Ventilation, and Air Conditioning (HVAC), but deploying MPC often requires substantial engineering effort. Reinforcement Learning (RL) may offer comparable performance with easier deployment, but its practical application for residential HVAC remains largely undemonstrated, leaving open questions related to occupant comfort and data requirements. To investigate these issues, we deployed one MPC variant and one model-based RL variant for one month each in an occupied house in a cold climate. The controllers adjusted an air-to-air heat pump's thermostat temperature setpoint based on measurements of the indoor temperature and the electric power used for heating. Relative to constant-setpoint operation, MPC saved 18.1\% (95\% confidence interval: 4.4 to 30.9\%) of weather-normalized heat pump energy and RL saved 20.9\% (2.6 to 38.3\%). MPC maintained acceptable occupant comfort. RL kept the house cooler, particularly during an initial adaptation phase, leading to three reports of occupant discomfort. The two algorithms had similar data requirements. We estimate that for a fresh deployment in another house, RL would take about one-third less engineering effort than MPC. While RL reduces deployment effort, it faces difficulties related to safe controller initialization and to mismatches between the modeled and true state and action spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。