优化跨数据中心大模型训练通信,最高提速64.62%
ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training

- 综合并行部署、调度与网络技术,全局优化跨域训练
- 实测相比生产配置提升64.62%,优于当前最优基线37.59%
- 适合大规模分布式训练系统设计者与算法工程师
大语言模型训练的快速扩展要求将GPU资源分布于多个数据中心乃至区域。我们称此范式为“跨域”训练。随着基础设施扩张,系统设计空间日益复杂,涵盖新型模型架构、硬件异构性及动态通信模式。基于Meta实际生产经验,我们揭示了在数个数据中心(含数十万张GPU)部署训练任务的复杂性。为加速大规模设计空间探索并支持前沿模型高效训练,我们深入分析三个关键设计维度:并行策略部署、并行调度机制与网络层技术。进而提出ScaleAcross Explorer,一个考虑各维度协同作用的优化器,实现跨域训练的整体优化。测试平台实验与仿真结果表明,该方案在多种设计点下相较生产配置最高提速64.62%,相较当前最优基线最高提速37.59%。
原文摘要 · Abstract (English)
The rapid scaling of large language model training requires distributing GPU resources across multiple data center buildings and regions. We refer to such paradigm as "scale-across" training. As infrastructure expands, the system design space becomes increasingly intricate, encompassing new model architectures, hardware heterogeneity, and evolving communication patterns. Drawing from Meta's production experience, we highlight the complexities of deploying training jobs across a few data centers housing hundreds of thousands of GPUs. To accelerate exploration of the large design space and to enable efficient training for frontier model development, we conduct in-depth characterization of three key design dimensions: parallelism placement, parallelism scheduling, and network layer technologies. We then propose ScaleAcross Explorer, an optimizer that considers the interplay of design dimensions and holistically optimizes scale-across training. Testbed experiments and simulations demonstrate up to 64.62% training speedups over production configuration and up to 37.59% training speedups over the state-of-the-art baseline across a wide range of design points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。