用强化学习优化风电场数据中心的能耗,提升风能利用率。
Toward an Energy-Optimized Operation of Data Centers Located in Wind Farms Using Reinforcement Learning

- 用强化学习在线调度计算任务,响应风电波动。
- 结合模仿学习与奖励塑形,使模型更充分使用早间免费风电。
- 在200天测试中表现接近最优解,适合风电融合数据中心场景。
本文研究强化学习作为在线控制器,在集成风力涡轮机的高性能计算(HPC)数据中心中实现可裁剪感知的任务迁移。提出一个可复现的固定日仿真框架,包含合成风速与电价信号及延迟完成反馈,便于扩展至复杂场景。以单风力涡轮机与单数据中心的最小场景为基准,发现纯强化学习存在明显的信用分配问题,倾向于低估早间免费风能。为此评估两种互补策略:基于优化的模仿学习与基于潜在的奖励塑形。在多种子训练和200天测试集上,近端策略优化(PPO)与带有额外在线更新的软演员-评论家(SAC)变体表现出优异性能,且两种策略均在相应配置下带来提升。与离线优化器相比仍存差距——后者拥有全天预见性,而强化学习仅依赖当前观测决策。该基准与消融实验为向多站点、连续时间场景扩展提供了透明基础。
原文摘要 · Abstract (English)
This paper studies Reinforcement Learning as an online controller for curtailment-aware workload shifting in wind-turbine-integrated high-performance computing (HPC) data centers. We introduce a reproducible fixed-day simulation framework with synthetic wind and price signals and delayed completion feedback, designed to be extensible toward more complex scenarios. As a controlled benchmarking basis, we then focus on the minimal case with one wind turbine and one co-located data center. In this setting, pure Reinforcement Learning exhibits a pronounced credit-assignment problem and tends to underuse free wind energy early in the day. We therefore evaluate two complementary countermeasures: optimization-based Imitation Learning and potential-based Reward Shaping. Across multi-seed training and a 200-day test set, Proximal Policy Optimization (PPO) and a Soft Actor-Critic (SAC) variant with an additional on-policy update routine achieve strong empirical performance among learned policies, and both Imitation Learning and Reward Shaping provide improvements in relevant configurations. A performance gap to the optimizer remains, which is expected: the optimizer plans offline with full-day foresight, whereas Reinforcement Learning must decide online from current observations without future realizations. The benchmark and ablation results provide a transparent basis for extending the approach toward richer multi-site and continuous-time scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。