用物理公式改进建筑能耗控制的奖励设计,让舒适度更真实、可解释。
PIRS: Physics-Informed Reward Shaping for SAC-Based Building Energy Management

- 用ISO 7730 PMV公式构建舒适度奖励,替代传统经验规则。
- 在50k步训练下,碳排放与电力消耗优于非物理基设计,峰值负荷降低1.78倍。
- 适合关注可解释性与标准对齐的智能建筑能源管理研究者。
occupants舒适度与电网友好型能效是相互冲突的目标,其联合优化高度依赖于深度强化学习(DRL)控制器中奖励函数的设计。然而,当前奖励设计仍多为随意设定:舒适度项常采用人工调参的启发式方法或简单的温度偏差代理,缺乏热舒适物理基础。本文提出PIRS(Physics-Informed Reward Shaping),将软演员-评论家(SAC)算法中的舒适度代理替换为基于ISO 7730预测平均投票(PMV)公式的加权多目标奖励。通过将舒适信号锚定在标准热舒适模型上,PIRS提升了奖励的可解释性,并提供了无须改动学习管道即可使用的标准基准舒适度代理。我们在CityLearn v2.1.2(2022挑战赛第一阶段)中评估了中央SAC智能体,训练50k步,五组随机种子。对比基准包括规则控制器(RBC)、手动设计奖励(E2)、仅能源奖励(E3)和简单温度偏差奖励(E4)。区域级关键绩效指标(KPI)以相对于RBC的比率报告:PIRS在成本、碳排放和电力方面表现与人工基线相当,显著优于非物理基设计,尤其在负荷爬坡(1.78倍优于~2.4倍RBC)和日峰值需求上。所有DRL策略在该训练预算下均优于RBC;我们坦诚承认此差距,并将PIRS定位为一种可解释、符合标准的奖励设计基础,而非在有限算力下对经典控制的全面超越。
原文摘要 · Abstract (English)
Occupant comfort and grid-aware energy efficiency are competing objectives whose joint optimization depends critically on how reward functions are specified in deep reinforcement learning (DRL) controllers for buildings. Yet reward design remains largely ad hoc: comfort terms are either hand-tuned heuristics or simple temperature-deviation proxies without explicit grounding in thermal-comfort physics. We present PIRS (Physics-Informed Reward Shaping), which replaces these ad-hoc comfort proxies with the ISO 7730 Predicted Mean Vote (PMV) formulation inside a weighted multi-objective reward for Soft Actor-Critic (SAC). By anchoring the comfort signal in the ISO 7730 PMV formulation, PIRS improves reward interpretability and provides a standards-grounded comfort proxy without changing any other component of the learning pipeline. We evaluate PIRS in CityLearn v2.1.2 (challenge 2022 phase 1) with a central SAC agent trained for 50k steps over five random seeds, and compare against a rule-based controller (RBC), a manually engineered reward (E2), an energy-only reward (E3), and a naive temperature-deviation comfort reward (E4). District-level key performance indicators (KPIs), reported as ratios versus RBC, show that PIRS attains cost, carbon, and electricity metrics on par with the manual baseline while substantially outperforming non-physics-grounded designs -- particularly on load ramping (1.78x vs. ~2.4x RBC) and daily peak demand. All DRL policies remain above RBC at this training budget; we interpret this gap honestly and position PIRS as an interpretable, standards-aligned foundation for reward design rather than a claim of dominance over classical control at limited compute.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。