为高密度AI机柜设计高效供电架构,避免电力浪费。
Designing Datacenter Power Delivery Hierarchies for the AI Era

- 构建多维度评估框架,融合真实部署与运维数据
- 2027年单机柜功率将达1MW,传统设计易导致电力闲置
- 关注可部署容量而非装机容量,适合数据中心长期规划
AI加速器需求激增,预计到2027年单个部署的机柜功耗密度将接近1MW。这给数据中心供电设计带来重大挑战。随着功率密度上升,针对不同目标密度设计的数据中心可能产生电力搁置,即无法使用其供电架构所预留的全部电力。设计需在数据中心全生命周期内保持高效,并适配多代硬件更新。由于电网电力资源在AI时代日益稀缺,电力利用率尤为重要。长期高效的供电架构设计困难,因机柜布局可行性、工作负载影响、成本等均受电气拓扑、部署粒度、放置策略、电力超分配和工作负载组合的共同影响。这些因素随时间演变,跨资源维度存在相互依赖,且难以进行闭式分析。为此,我们开发了一个评估框架,基于真实中服务器部署、超分配及退役序列,综合吞吐量、功耗与成本指标。该框架结合了微软Azure生产数据中的GPU、计算与存储部署预测模型。结果表明,多资源搁置显著影响可部署容量、实际资本支出与交付性能;并量化了从机柜级到机架级的AI系统功率密度上升对上述结果的影响。对于AI数据中心设计,关键规划目标不是装机兆瓦数,而是随时间变化的可部署容量。
原文摘要 · Abstract (English)
Demand for AI accelerators is rapidly increasing rack power density, with projections approaching 1MW per deployment by 2027. This poses a major challenge for datacenter power delivery designers. As power densities increase, a datacenter designed for a different target density may strand power, i.e., may be unable to use all the power that its delivery hierarchy has provisioned. Designs must remain efficient over long datacenter lifetimes and multiple hardware generations. Power utilization is particularly important as grid power capacity is a scarce resource in the AI era. Designing an efficient power delivery hierarchy for the long run is difficult because rack placement feasibility, workload impact, and cost depend jointly on electrical topology, deployment granularity, placement policy, power oversubscription, and workload mix. Moreover, each of these factors evolve over time, have inter-dependencies across multiple resource dimensions, and generally do not lend themselves to closed-form analysis. To address this challenge, we develop a framework for evaluating datacenter power delivery designs using throughput, power, and cost metrics over realistic arrival, oversubscription, and decommissioning sequences. The framework combines projection models for GPU, compute, and storage deployments with operational factors grounded in production data from Microsoft Azure. Our results show that multi-resource stranding materially changes deployable capacity, effective capital expenditure, and delivered performance, and quantify how rising density from rack- and pod-scale AI systems shapes these outcomes. For AI datacenter design, the relevant planning objective is not installed megawatts, but deployable capacity over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。