解决MoE模型调度中的指数级数据汇聚问题,提升网络利用率
Incast-Free MoE Rate-Based Scheduling

- 提出基于速率的主动公平调度框架,避免传统轮询导致的数据汇聚
- 实测可消除数据汇聚,实现近100%链路利用率,降低集体完成时间
- 适合大规模MoE训练系统部署,尤其关注网络性能优化的研究者
混合专家(MoE)架构已成为大语言模型的核心;然而,其典型的轮询(RR)调度方式引入了显著瓶颈。本文首次揭示了RR调度在MoE流量下会引发指数级的入聚集(incast)现象。为此,我们提出一种专为MoE工作负载设计的主动公平调度框架,有效防止网络过载。同时,阐明了该方案在网卡(NIC)中的实现路径。通过真实与合成工作负载的广泛仿真验证,该框架能持续消除入聚集,维持接近100%的链路利用率,并显著降低集体完成时间(CCT)。
原文摘要 · Abstract (English)
Mixture of Experts (MoE) architectures have become key to large language models; however, their typical round-robin (RR) scheduling introduces significant bottlenecks. In this paper, we demonstrate that RR causes a previously-undiscovered exponential incast phenomenon with MoE traffic. We propose an alternative proactive fair scheduling framework tailored for MoE workloads, which effectively prevents fabric oversubscription. We also outline how it can be implemented in NICs. Finally, through extensive simulations with real and synthetic workloads, we demonstrate that this framework consistently eliminates incast, maintains a near-100% link utilization, and reduces Collective Completion Time (CCT).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。