用深度强化学习优化数据中心能源调度,降本减碳且保障服务稳定
Green Energy Management for Sustainable Data Centers Using Deep Reinforcement Learning
- 基于PPO与时空注意力网络的DRL框架,动态协调光伏、风电等多源能源
- 能耗降低38%,碳排放和停机风险显著下降,能效达83.7%
- 适合关注绿色计算、智能调度的科研与工程人员
数字服务的指数级增长使数据中心成为现代经济中最耗能的基础设施之一,带来运营成本、碳排放及可再生能源融合的严峻挑战。本文提出一种基于深度强化学习(DRL)的数据中心能源管理框架,旨在高随机性条件下动态协调光伏发电、风力发电、电池储能与传统电网电力。该框架将能源管理建模为马尔可夫决策过程,采用增强型近端策略优化(PPO)代理,结合混合长短期记忆与时间注意力架构,精准捕捉负载动态与可再生能源波动。设计多目标奖励函数,联合最小化能源成本、碳排放与服务等级协议(SLA)违规,同时提升储能利用率。在三个数据集上的实验表明,该框架相比规则启发式方法降低38%能耗,优于最强DRL基线4.6%,且保持1.5%的低SLA违规率和83.7%的能源效率。消融实验验证各组件贡献,超参数敏感性分析证明方法在多种配置下的鲁棒性。
原文摘要 · Abstract (English)
The exponential growth of digital services has positioned data centers among the most energy-intensive infrastructures in the modern economy, raising critical concerns regarding operational costs, carbon emissions, and the sustainable integration of renewable energy sources. This paper proposes a novel Deep Reinforcement Learning (DRL)-based energy management framework for data centers, designed to dynamically coordinate solar photovoltaic generation, wind power, battery storage systems, and conventional grid electricity under highly stochastic operational conditions. The proposed framework formulates the energy management problem as a Markov Decision Process and employs a Proximal Policy Optimization (PPO) agent augmented with a hybrid Long Short-Term Memory and temporal attention architecture, enabling accurate modeling of workload dynamics and renewable generation variability. A multi-objective reward function jointly minimizes energy costs, carbon emissions, and service-level agreement (SLA) violations while promoting efficient storage utilization. Extensive experiments conducted on three datasets demonstrate that the proposed framework achieves a 38\% reduction in energy costs compared to rule-based heuristics and outperforms the strongest DRL baseline by 4.6\%, while maintaining an SLA violation rate as low as 1.5\% and an energy efficiency of 83.7\%. Ablation studies confirm the individual contribution of each architectural component, and hyperparameter sensitivity analysis validates the robustness of the approach across a range of configurations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。