用离线强化学习优化数据中心冷却,节能14%~21%且不违反安全约束
Data Center Cooling System Optimization Using Offline Reinforcement Learning
- 设计基于图神经网络的物理感知模型,捕捉机房动态与物理规律
- 仅用真实运行数据实现样本高效、鲁棒的策略学习,实测节能14%~21%
- 适合工业界解决数据少、安全要求高的控制优化问题
信息技术与人工智能的快速发展推动了全球数据中心(DC)行业的迅猛扩张,随之带来巨大的电力消耗。在典型数据中心中,冷却系统耗电占比达30%~40%,远超服务器本身,亟需新型节能优化技术。然而,优化此类工业系统面临诸多挑战:缺乏可靠仿真环境、历史数据有限,以及对安全性与控制鲁棒性的严苛要求。本文提出一种新型物理信息引导的离线强化学习(RL)框架,用于数据中心冷却系统的能效优化。该框架采用定制的图神经网络架构,建模机房内部复杂的动态模式与物理依赖关系,并遵守基本的时间反演对称性。得益于其良好行为且可泛化的状态-动作表示,模型可在有限的真实运行数据下实现高效、鲁棒的隐空间离线策略学习。该框架已在大型生产级数据中心成功部署并验证,用于闭环控制空气冷却单元(ACUs)。我们在生产环境中进行了总计2000小时的短期与长期实验,结果表明,该方法在不违反任何安全或运行约束的前提下,使数据中心冷却系统实现14%~21%的能源节省。结果证明了离线强化学习在解决广泛的数据受限、安全关键型工业控制问题上具有巨大潜力。
原文摘要 · Abstract (English)
The recent advances in information technology and artificial intelligence have fueled a rapid expansion of the data center (DC) industry worldwide, accompanied by an immense appetite for electricity to power the DCs. In a typical DC, around 30~40% of the energy is spent on the cooling system rather than on computer servers, posing a pressing need for developing new energy-saving optimization technologies for DC cooling systems. However, optimizing such real-world industrial systems faces numerous challenges, including but not limited to a lack of reliable simulation environments, limited historical data, and stringent safety and control robustness requirements. In this work, we present a novel physics-informed offline reinforcement learning (RL) framework for energy efficiency optimization of DC cooling systems. The proposed framework models the complex dynamical patterns and physical dependencies inside a server room using a purposely designed graph neural network architecture that is compliant with the fundamental time-reversal symmetry. Because of its well-behaved and generalizable state-action representations, the model enables sample-efficient and robust latent space offline policy learning using limited real-world operational data. Our framework has been successfully deployed and verified in a large-scale production DC for closed-loop control of its air-cooling units (ACUs). We conducted a total of 2000 hours of short and long-term experiments in the production DC environment. The results show that our method achieves 14~21% energy savings in the DC cooling system, without any violation of the safety or operational constraints. Our results have demonstrated the significant potential of offline RL in solving a broad range of data-limited, safety-critical real-world industrial control problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。