arXiv:2511.00116cs.LGcs.AI2025-11NeurIPS被引 5

用强化学习优化数据中心液冷系统,提升能效与可靠性。

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

  • 构建高保真数字孪生环境,支持端到端液冷控制优化。
  • 实现多目标实时优化,降低能耗同时保障散热安全。
  • 支持可解释性控制与大模型智能代理,适合研究者与运维人员使用。

液冷是应对高密度数据中心日益增长的AI算力负载下热管理的关键技术。为实现更高能效与可靠性,机器学习控制器至关重要。本文提出LC-Opt,一个面向可持续液冷(LC)的基准环境,用于高性能计算(HPC)系统中强化学习(RL)控制策略的研究。基于美国橡树岭国家实验室前沿超算系统的高保真数字孪生基础,LC-Opt提供从冷却塔到机柜、服务器刀片组的全流程Modelica建模,覆盖站点级至设备级。通过Gymnasium接口,RL智能体可优化液冷供应温度、流量、机柜级阀门动作以及冷却塔设定点,在动态工作负载条件下实现控制。该环境构建了兼顾局部温控与全局能效的多目标实时优化挑战,并支持热回收单元(HRU)等扩展组件。我们对比了集中式与分布式多智能体RL方法,展示了将策略蒸馏为决策树和回归树以实现可解释控制,还探索了基于大语言模型的智能代理架构,通过自然语言解释控制行为,增强用户信任并简化管理。LC-Opt推动了精细化、可定制液冷模型的开放获取,助力机器学习社区、运营商与厂商开发可持续的数据中心液冷解决方案。

原文摘要 · Abstract (English)

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

液冷优化强化学习数字孪生能效提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。