用AI自动生成数据中心调控策略,应对动态负载与协议变化。
DCoPilot: Generative AI-Empowered Policy Adaptation for Dynamic Data Center Operations
- 结合大模型与超网络,实现奖励函数与策略权重的联合生成。
- 在五类任务中近零违规,性能超越所有基线方法。
- 适合需要快速响应配置变更的数据中心运维团队。
现代数据中心搭载人工智能专用设备,功耗密度高且工作负载波动剧烈,分钟级动态调控对安全节能至关重要。然而,人工设计分段深度强化学习(DRL)代理难以跟上频繁的动态变化与服务级别协议(SLA)调整,导致控制策略滞后,可能引发服务中断。为此,本文提出DCoPilot,一种面向动态数据中心运行的生成式控制策略混合框架。该框架融合两种生成范式:大语言模型(LLM)生成结构化奖励形式,超网络(hypernetwork)生成策略权重。系统通过三个协同阶段运作:(i) 模拟规模扩展,在多样化的仿真就绪(SimReady)场景中压力测试奖励候选;(ii) 元策略蒸馏,训练超网络根据SLA与场景嵌入输出策略权重;(iii) 在线适应,实现对更新后规范的零样本策略生成。在涵盖多种数据中心组件的五类控制任务中评估,DCoPilot实现近零约束违反,且在不同规范变化下均优于所有基线。消融实验验证了基于LLM的统一奖励生成在促进超网络稳定收敛中的有效性。
原文摘要 · Abstract (English)
Modern data centers (DCs) hosting artificial intelligence (AI)-dedicated devices operate at high power densities with rapidly varying workloads, making minute-level adaptation essential for safe and energy-efficient operation. However, manually designing piecewise deep reinforcement learning (DRL) agents cannot keep pace with frequent dynamics shifts and service-level agreement (SLA) changes of an evolving DC. This specification-to-policy lag causes a lack of timely, effective control policies, which may lead to service outages. To bridge the gap, we present DCoPilot, a hybrid framework for generative control policies in dynamic DC operation. DCoPilot synergizes two distinct generative paradigms, i.e., a large language model (LLM) that performs symbolic generation of structured reward forms, and a hypernetwork that conducts parametric generation of policy weights. DCoPilot operates through three coordinated phases: (i) simulation scale-up, which stress-tests reward candidates across diverse simulation-ready (SimReady) scenes; (ii) meta policy distillation, where a hypernetwork is trained to output policy weights conditioned on SLA and scene embeddings; and (iii) online adaptation, enabling zero-shot policy generation in response to updated specifications. Evaluated across five control task families spanning diverse DC components, DCoPilot achieves near-zero constraint violations and outperforms all baselines across specification variations. Ablation studies validate the effectiveness of LLM-based unified reward generation in enabling stable hypernetwork convergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。