arXiv:2608.09453cs.LGcs.AI2026-08中稿 · UKCI 2026, Will be…

用强化学习让热泵不启停,实现更耐用的连续调节控制

Learning to Modulate, Not to Cycle: Soft Actor---Critic Recovers Inverter-Style Heat-Pump Control

论文配图:Learning to Modulate, Not to Cycle: Soft Actor---Critic Recovers Inverter-Style Heat-Pump Control
图 1 · 摘自论文原文
  • 在奖励中加入压缩机损耗项,引导算法学习持续调节而非频繁启停
  • SAC算法实现零启停(0次/天),比基线减少90.7%热不适
  • 适合关注设备寿命与舒适性平衡的智能建筑控制研究者

住宅热泵的压缩机启停是导致磨损的主要原因,但现有强化学习(RL)控制器通常只优化能耗与热舒适性,忽略启停频率。本文在控制奖励中引入压缩机损耗的平准化项,研究不同RL算法的行为差异。在BOPTEST最佳测试水力热泵案例的相同马尔可夫决策过程中训练Soft Actor-Critic(SAC)和近端策略优化(PPO),发现SAC学习到一种持续调节策略,使压缩机永久运行——即变频热泵的工作原理,实现每天0次启停;而PPO退化为开关控制,启停次数超过基线。在BOPTEST模拟器上,SAC策略使热不适降低高达90.7%,成本仅增加11.5%,同时完全消除基线启停。

原文摘要 · Abstract (English)

On--off cycling is the main cause of compressor wear in residential heat pumps, yet reinforcement learning (RL) controllers for buildings typically optimise only energy cost and thermal comfort, ignoring how much the learned policy cycles. We add a levelised compressor-wear term to the control reward and study how the resulting behaviour depends on the RL algorithm. Training Soft Actor---Critic (SAC) and Proximal Policy Optimisation (PPO) on an identical Markov decision process for the BOPTEST bestest hydronic heat pump case, we find that SAC learns a continuous modulation policy that keeps the compressor permanently engaged---the operating principle of an inverter-driven heat pump---achieving zero start-ups per day, whereas PPO collapses to bang-bang control that cycles more than the baseline. On the BOPTEST emulator the SAC policy cuts thermal discomfort by up to 90.7% for an 11.5% cost increase, while eliminating all baseline cycling.

强化学习热泵控制设备寿命连续调节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。