arXiv:2605.28168cs.AI2026-05中稿 · OccuSys 2026, co-l…

用大模型优化建筑能耗控制,让不同人群更舒适且节能

OccuReward: LLM-Guided Occupant-Centric Reward Shaping for Demographic Equity in Grid-Interactive Buildings

论文配图:OccuReward: LLM-Guided Occupant-Centric Reward Shaping for Demographic Equity in Grid-Interactive Buildings
图 1 · 摘自论文原文
  • 用大模型迭代生成兼顾公平的奖励函数
  • 老年女性满意度提升567%,能耗降3.2%
  • 适合关注智能建筑公平性的研究者

大型语言模型(LLMs)在生成深度强化学习(DRL)建筑能源管理奖励函数方面展现出潜力,但其对异质人口群体在居住者舒适度上的潜在差异影响尚未被探索。本文提出OccuReward框架,研究大模型驱动的奖励设计如何影响人口公平性。贡献包括:引入舒适度公平指数(CEI)作为新反馈信号;提出一种迭代式、公平感知的LLM奖励塑造方法;以及对优化后目标下DRL代理性能的分析。基于ASHRAE全球热舒适数据库II的四个实证居民画像(13,440个投票),我们在CityLearn v2中部署了软演员-评论家(Soft Actor-Critic)智能体。通过Gemini API生成奖励函数逻辑与权重,而非每步推理,共进行三轮优化。15次实验显示,初始阶段老年女性满意度最低。至第三轮,公平感知的LLM优化激活特定奖励组件,使年轻男性满意度提升17.6%,中年女性提升28.2%,健康敏感群体提升53.8%,老年女性提升567%,同时能耗降低3.2%。结果表明,奖励干预显著改善公平性,但算法驱动控制器中的群体差异仍存,需进一步研究建筑系统中的算法公平性。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated promising capability in generating reward functions for deep reinforcement learning (DRL)-based building energy management. However, their potential to exhibit or exacerbate disparities in occupant comfort across heterogeneous demographic populations remains unexplored. We present OccuReward, a framework investigating how LLM-mediated reward design affects demographic equity. Our contribution is three-fold: the introduction of the Comfort Equity Index (CEI) as a novel feedback signal; a methodology for iterative, equity-aware LLM reward shaping; and a performance analysis of DRL agents under these refined objectives. Utilizing four empirically grounded occupant profiles from the ASHRAE Global Thermal Comfort Database II (13,440 votes), we deploy a Soft Actor-Critic agent in CityLearn v2. Our approach employs the Gemini API to generate reward function logic and weights--rather than performing per-step inference--across three refinement rounds. Results across 15 experimental runs reveal that elderly female occupants consistently experience the lowest satisfaction in initial rounds. By Round 3, equity-aware LLM refinement activates specific reward components that improve satisfaction for Young Males (+17.6%), Mid-aged Females (+28.2%), Health Sensitive (+53.8%), and Elderly Females (+567%), while simultaneously reducing energy costs by 3.2%. Our findings highlight that while reward-level intervention significantly improves equity, demographic disparities in AI-driven controllers persist, necessitating further research into algorithmic fairness in building systems.

建筑能源公平性大模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。