为动态通信的专家模型训练设计可实时重构的光电混合网络
MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training
- 基于光路交换实现局部动态拓扑重构,适应专家模型的运行时通信变化
- 在32张A100 GPU上实测,100Gbps下训练效率提升1.2-1.5倍,400Gbps下达1.9-2.3倍
- 适合大规模专家模型训练系统设计者与高性能计算架构研究者
Mixture-of-Expert (MoE) 模型通过按令牌选择性激活不同子网络(专家)来超越传统模型。这种门控计算产生动态通信模式,无法预先确定,挑战了现有训练中保持静态的GPU互连。本文提出首个此类系统 MixNet,支持分布式 MoE 训练过程中的拓扑重构。基于生产环境测量发现 MoE 通信具有强局部性,降低对全局重构的需求。在此基础上,我们利用光路交换(OCS)在现有电互连上构建区域可重构的高带宽域,兼顾可扩展性与快速适应性。我们使用通用硬件搭建了完整 MixNet 原型,并配备定制集体通信运行时,在32张A100 GPU上实现了状态前沿MoE模型的训练与运行时拓扑重构。大规模包级仿真表明,MixNet 在100 Gbps和400 Gbps链路带宽下,相比非阻塞胖树结构,分别将四类代表性MoE模型的训练成本效率提升1.2-1.5倍和1.9-2.3倍。
原文摘要 · Abstract (English)
Mixture-of-Expert (MoE) models outperform conventional models by selectively activating different subnets, named experts, on a per-token basis. This gated computation generates dynamic communications that cannot be determined beforehand, challenging the existing GPU interconnects that remain static during the distributed training process. In this paper, we advocate for a first-of-its-kind system, called MixNet, that unlocks topology reconfiguration during distributed MoE training. Towards this vision, we first perform a production measurement study and show that the MoE dynamic communication pattern has strong locality, alleviating the requirement of global reconfiguration. Based on this, we design and implement a regionally reconfigurable high-bandwidth domain on top of existing electrical interconnects using optical circuit switching (OCS), achieving scalability while maintaining rapid adaptability. We have built a fully functional MixNet prototype with commodity hardware and a customized collective communication runtime that trains state-of-the-art MoE models with in-training topology reconfiguration across 32 A100 GPUs. Large-scale packet-level simulations show that MixNet delivers comparable performance as the non-blocking fat-tree fabric while boosting the training cost efficiency (e.g., performance per dollar) of four representative MoE models by 1.2x-1.5x and 1.9x-2.3x at 100 Gbps and 400 Gbps link bandwidths, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。