用三维光互连加速万亿参数MoE模型训练,提速2.7倍。
Accelerating Frontier MoE Training with 3D Integrated Optics
- 采用3D堆叠光学技术连接数百个GPU,突破电互连距离限制。
- 带宽和端口数提升8倍,实现超大规模模型高效训练。
- 适合超大规模AI训练团队与数据中心架构设计者参考。
AI算力需求持续增长,推动计算、内存与互连性能的协同进步。随着传统半导体缩放放缓,高速互连成为新的扩展引擎,通过将多个GPU连接为单一低延迟、高带宽的计算域,实现更大逻辑GPU。早期扩展架构依赖铜线互连以降低成本和功耗,但被动电互连的最大传输距离约1米,使扩展域受限于单机柜。3D堆叠光电子与逻辑技术提供了可扩展、低功耗的解决方案,支持跨多个机柜连接数百个GPU封装(数千个GPU)。本文研究扩展技术的设计权衡,表明前沿大模型需要新型光子方案以达成激进的能效与性能目标。我们建模了3D CPO(Passage)赋能的GPU与交换机在扩展域中的表现,用于训练超过一万亿参数的前沿混合专家(MoE)模型。结果表明,3D CPO带来的带宽与端口数显著提升,使扩展能力提高8倍,为扩展域内提供多维并行新可能,最终实现训练时间缩短2.7倍,开启前所未有的模型规模扩展空间。
原文摘要 · Abstract (English)
The unabated growth in AI workload demands is driving the need for concerted advances in compute, memory, and interconnect performance. As traditional semiconductor scaling slows, high-speed interconnects have emerged as the new scaling engine, enabling the creation of larger logical GPUs by linking many GPUs into a single, low-latency, high-bandwidth compute domain. While initial scale-up fabrics leveraged copper interconnects for their power and cost advantages, the maximum reach of passive electrical interconnects (approximately 1 meter) effectively limits the scale-up domain to within a single rack. The advent of 3D-stacked optics and logic offers a transformative, power-efficient scale-up solution for connecting hundreds of GPU packages (thousands of GPUs) across multiple data center racks. This work explores the design tradeoffs of scale-up technologies and demonstrates how frontier LLMs necessitate novel photonic solutions to achieve aggressive power and performance targets. We model the benefits of 3D CPO (Passage) enabled GPUs and switches within the scale-up domain when training Frontier Mixture of Experts (MoE) models exceeding one trillion parameters. Our results show that the substantial increases in bandwidth and radix enabled by 3D CPO allow for an 8X increase in scale-up capability. This affords new opportunities for multi-dimensional parallelism within the scale-up domain and results in a 2.7X reduction in time-to-train, unlocking unprecedented model scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。