为万亿参数大模型设计高效数据中心,提出全连接网络与系统协同优化方案。
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
- 提出FullFlat全连接光网络,实现节点间低延迟高带宽互联。
- 通过硬件加速通信与扩大扩展域,提升模型吞吐量与计算利用率。
- 适配稀疏与密集模型,为下一代AI数据中心提供可落地的设计指南。
大型语言模型(如参数达1.8万亿的GPT-4)的迅猛发展,要求对数据中心架构进行根本性重构,以保障可扩展性、效率和成本效益。本文提出一种联合设计框架,综合评估计算能力(FLOPS)、HBM带宽与容量、双层与FullFlat光网络拓扑、扩展域规模,以及主流并行与优化策略。引入并评估FullFlat网络架构,该架构在所有节点间提供一致的高带宽、低延迟连接,显著提升性能与可扩展性。通过敏感性分析,量化了计算与通信重叠、硬件加速集合操作、扩大扩展域及增加内存容量带来的收益。研究涵盖稀疏(专家混合)与密集变压器架构的LLM,揭示系统设计对模型每秒浮点运算利用率(MFU)与整体吞吐量的影响。采用解析性能建模工具,预测结果与真实测量值偏差小于10%。研究成果为支持万亿参数模型的数据中心设计提供了可操作的洞察与实用路线图。
原文摘要 · Abstract (English)
The explosive growth of Large Language Models (LLMs), such as GPT-4 with 1.8 trillion parameters, demands a fundamental rethinking of data center architecture to ensure scalability, efficiency, and cost-effectiveness. Our work provides a comprehensive co-design framework that jointly explores FLOPS, HBM bandwidth and capacity, multiple network topologies (two-tier vs. FullFlat optical), the size of the scale-out domain, and popular parallelism/optimization strategies used in LLMs. We introduce and evaluate FullFlat network architectures, which provide uniform high-bandwidth, low-latency connectivity between all nodes, and demonstrate their transformative impact on performance and scalability. Through detailed sensitivity analyses, we quantify the benefits of overlapping compute and communication, leveraging hardware-accelerated collectives, widening the scale-out domain, and increasing memory capacity. Our study spans both sparse (mixture of experts) and dense transformer-based LLMs, revealing how system design choices affect Model FLOPS Utilization (MFU = Model FLOPS per token * Observed tokens per second / Peak FLOPS of the hardware) and overall throughput. For the co-design study, we utilized an analytical performance modeling tool capable of predicting LLM runtime within 10% of real-world measurements. Our findings offer actionable insights and a practical roadmap for designing AI data centers that can efficiently support trillion-parameter models, reduce optimization complexity, and sustain the rapid evolution of AI capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。