构建可评估变化的建筑环境基准,提升多目标强化学习泛化能力
BEAVER: Building Environments with Assessable Variation for Evaluating Multi-Objective Reinforcement Learning
- 提出多目标情境强化学习框架,显式建模气候与热传导等上下文变量
- 新基准显示现有方法在环境变化下性能下降,凸显动态信息重要性
- 适合关注建筑能源控制、多目标鲁棒性评估的研究者
近年来,基于强化学习(RL)的建筑能源管理代理取得显著进展。尽管在模拟或受控环境中表现良好,但其在效率和跨建筑动态及运行场景的泛化能力仍不明确。本文正式刻画了跨环境、多目标建筑能源管理的任务泛化空间,提出多目标情境强化学习问题。该框架有助于理解策略在不同气候条件和热对流动态下,于舒适度与能耗等多重目标间转移的挑战。我们提供了一种在真实建筑RL环境中参数化上下文信息的系统方法,并构建了一个新型基准,用于评估实际建筑控制任务中可泛化的强化学习算法。结果表明,现有方法虽能实现合理的目标权衡,但在特定环境变化下性能下降,强调了在策略学习中融入依赖动态的上下文信息的重要性。
原文摘要 · Abstract (English)
Recent years have seen significant advancements in designing reinforcement learning (RL)-based agents for building energy management. While individual success is observed in simulated or controlled environments, the scalability of RL approaches in terms of efficiency and generalization across building dynamics and operational scenarios remains an open question. In this work, we formally characterize the generalization space for the cross-environment, multi-objective building energy management task, and formulate the multi-objective contextual RL problem. Such a formulation helps understand the challenges of transferring learned policies across varied operational contexts such as climate and heat convection dynamics under multiple control objectives such as comfort level and energy consumption. We provide a principled way to parameterize such contextual information in realistic building RL environments, and construct a novel benchmark to facilitate the evaluation of generalizable RL algorithms in practical building control tasks. Our results show that existing multi-objective RL methods are capable of achieving reasonable trade-offs between conflicting objectives. However, their performance degrades under certain environment variations, underscoring the importance of incorporating dynamics-dependent contextual information into the policy learning process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。