arXiv:2605.04074cs.LGcs.AI2026-05KDD

用物理模型提升AI数据中心GPU功耗预测精度,提前5-80分钟预警电力波动。

A Physics-Aware Framework for Short-Term GPU Power Forecasting of AI Data Centers

论文配图:A Physics-Aware Framework for Short-Term GPU Power Forecasting of AI Data Centers
图 1 · 摘自论文原文
  • 融合热力学物理规律的DLinear模型,建模GPU功耗与温度、算力、内存的动态关系。
  • 在真实数据集上,误差比顶尖模型低0.78%至51.82%,尤其在负载突变时仍保持物理一致性。
  • 适合关注数据中心能效优化、电力调度的工程师和研究者参考。

AI数据中心因计算任务异构性导致功耗快速波动,例如大语言模型的推理与训练功耗差异显著,可能引发电网不稳定。本文首次提出物理感知的DLinear时间序列模型(PI-DLinear),可准确预测未来5至80分钟的AI数据中心功耗。该模型基于多节点集中式热阻容(RC)网络,结合牛顿冷却定律,通过新推导的时间依赖常微分方程(ODE),分别建模并关联GPU计算、内存利用率与温度对功耗的影响。模型在真实数据中心数据集上训练与评估,不仅整体精度优于当前最先进(SOTA)的基于Transformer与非Transformer模型,且在功率降频和负载瞬态事件中仍遵循底层物理规律。相比SOTA模型,其均方误差(MSE)降低0.782%~39.08%,平均绝对误差(MAE)降低0.993%~51.82%,均方根误差(RMSE)降低0.370%~22.28%。

原文摘要 · Abstract (English)

AI data centers experience rapid fluctuations in power demand due to the heterogeneity of computational tasks that they have to support. For example, the power profile of inference and training of large language models (LLMs) is quite distinct and big divergences can result in the instability of the underlying electricity grid. In this paper we propose, to the best of our knowledge, the first physics-informed DLinear time-series model that can accurately forecast power utilization of an AI data center 5-80 minutes (short-term forecasting) into the future. The physics, based on a multi-node lumped thermal resistance-capacitance (RC) network consistent with Newton's law of cooling, is captured using newly derived time-dependent ordinary differential equations (ODE) that separately models and interlinks power consumption with the GPU compute and memory utilization and temperature. The resulting model, that we refer to as PI-DLinear, trained and evaluated on a real AI data center dataset and is not only more accurate than the state-of-the-art (SOTA) models tested, but the forecast profile respects the underlying physics under power throttling and load transient events. Relative to the SOTA transformer-based and non-transformer-based models, improvements in forecasting accuracy (averaged across all look-back and prediction windows) range from 0.782%-39.08% for MSE, 0.993%-51.82% for MAE, and 0.370%-22.28% for RMSE.

功耗预测物理模型GPU数据中心

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。