针对AI负载突增,提出可提前预警容量压力的轻量级预测框架。
Learning Burst-Aware Early Warning Models for Capacity Stress under AI Workload Surges in Hyperscale Data Centers

- 融合负载强度、时间变化与系统压力信号,用轻量树模型捕捉非线性关系。
- 在真实场景模拟下,召回率达0.914,AUC达0.697,显著优于基线。
- 适合需主动调控资源的超大规模数据中心运维团队使用。
大规模AI工作负载(尤其是大语言模型训练与推理)正深刻改变超大规模数据中心的运行模式。与传统云工作负载不同,AI任务具有突发性、高密集和快速变化的资源需求,常引发突发容量压力,现有基于阈值的被动响应机制难以应对。本文提出一种面向部署的、具备突发感知能力的早期预警框架,用于在容量压力发生前进行主动预测。将问题建模为多变量遥测窗口上的高召回率预测任务,明确目标是在系统退化前触发操作干预。框架整合负载强度、时间变化和系统压力信号,采用轻量级树模型捕捉高度不平衡环境中的非线性交互。为评估系统在真实条件下的表现,引入一种AI负载突增注入方法,模拟大规模AI系统中观测到的突发需求模式。基于XGBoost的模型在测试中达到ROC AUC 0.697和平均精度(AP)0.670,显著优于基线方法。在面向部署的阈值选择下,框架实现0.914的召回率,可在可接受误报代价下检测绝大多数压力期。此外,结果表明该框架可集成至运维控制环路,支持如工作负载限速与资源扩容等主动措施。研究凸显了高召回率学习型预警系统在AI驱动时代保障数据中心弹性与自适应运行的实际价值。
原文摘要 · Abstract (English)
The rapid growth of large-scale AI workloads, particularly Large Language Model (LLM) training and inference, is fundamentally reshaping the operational dynamics of hyperscale data centers. Unlike traditional cloud workloads, AI-driven jobs exhibit bursty, high-intensity, and rapidly shifting resource demands, often leading to sudden capacity stress that cannot be effectively handled by reactive threshold-based mechanisms. In this paper, we propose a deployment-oriented, burst-aware early warning framework for proactive capacity stress prediction under AI workload surges. We formulate the problem as a high-recall forecasting task over multivariate telemetry windows, with the explicit goal of enabling operational intervention before system degradation occurs. The proposed framework integrates workload intensity, temporal variation, and system pressure signals, and employs a lightweight tree-based learning model to capture nonlinear interactions in highly imbalanced environments. To evaluate the system under realistic conditions, we introduce an AI workload surge injection methodology that simulates burst-driven demand patterns observed in large-scale AI systems. Our XGBoost-based model achieves an ROC AUC of 0.697 and an AP of 0.670, significantly outperforming baseline methods. Under deployment-oriented threshold selection, the framework achieves a Recall of 0.914, enabling the detection of the majority of stress-prone periods with acceptable false-alarm cost. Beyond predictive performance, we show how the proposed framework can be integrated into operational control loops to support proactive actions such as workload throttling and resource scaling. Our results highlight the practical value of high-recall, learning-based early warning systems in enabling resilient and adaptive data center operations in the era of AI-driven workloads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。