解决大规模AI训练中因算力波动导致的电力不稳问题
Power Stabilization for AI Training Datacenters
- 从软件、芯片到数据中心多层协同优化供电
- 实测显示可将电力波动幅度降低超60%
- 适合超大规模AI训练平台的工程团队参考
数千张GPU并行的大型AI训练任务带来独特的电源管理挑战。由于训练过程中计算与通信阶段交替进行,算力密集期功耗远高于通信期,导致显著的功率波动。随着训练规模扩大,这种波动幅度持续上升,其频率谱若与电网固有频率共振,可能造成物理损坏。为保障大规模训练安全扩容,本文基于真实生产数据,探索了从软件、GPU硬件到数据中心基础设施的创新解决方案,评估各方法优劣,并提出综合应对策略。通过真实硬件与微软自研云电源模拟器联合测试,验证了干预措施在实际场景中的有效性。
原文摘要 · Abstract (English)
Large Artificial Intelligence (AI) training workloads spanning several tens of thousands of GPUs present unique power management challenges. These arise due to the high variability in power consumption during the training. Given the synchronous nature of these jobs, during every iteration there is a computation-heavy phase, where each GPU works on the local data, and a communication-heavy phase where all the GPUs synchronize on the data. Because compute-heavy phases require much more power than communication phases, large power swings occur. The amplitude of these power swings is ever increasing with the increase in the size of training jobs. An even bigger challenge arises from the frequency spectrum of these power swings which, if harmonized with critical frequencies of utilities, can cause physical damage to the power grid infrastructure. Therefore, to continue scaling AI training workloads safely, we need to stabilize the power of such workloads. This paper introduces the challenge with production data and explores innovative solutions across the stack: software, GPU hardware, and datacenter infrastructure. We present the pros and cons of each of these approaches and finally present a multi-pronged approach to solving the challenge. The proposed solutions are rigorously tested using a combination of real hardware and Microsoft's in-house cloud power simulator, providing critical insights into the efficacy of these interventions under real-world conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。