arXiv:2608.26453cs.LGcs.DC2026-08

让网络主动参与分布式训练,提升跨广域网训练效率。

Distributed Training using an Intelligent Network

  • 利用组播和FPGA聚合网络流量,缓解跨区域通信瓶颈。
  • 基于网络拓扑设计旋转集群同步策略,最大化信息交换效率。
  • 在九城真实广域网实验中验证效果,接近本地训练性能。

跨广域网(WAN)的分布式训练面临带宽有限、延迟高和拓扑不均的挑战。本文提出让网络成为训练过程的主动参与者:系统层面采用组播技术复制出站流量,并利用线内FPGA聚合入站流量,缓解进出带宽瓶颈;算法层面构建优化框架,生成基于网络拓扑与技术特性的丰富同步调度方案(如旋转子群),以最大化信息交换。在模拟九城市、基于DoubleZero可编程广域网的真实场景中验证,最优调度随网络能力动态调整,显著缩小与本地共置训练的性能差距。

原文摘要 · Abstract (English)

Distributed training across a wide area network (WAN) is challenging, as continuous parameter exchange by islands of compute is constrained by limited bandwidth, high latency, and uneven topology. We propose making the network an active participant in training. On the systems side, such networks should leverage (i) multicast technology to replicate outbound traffic and (ii) in-line FPGAs to aggregate inbound traffic, to ease egress and ingress bottlenecks. These technologies are used for training across workers within a data center, but this paper extends them to the WAN. On the algorithms side, we develop an optimization framework that produces rich synchronization schedules (namely, rotating cliques of islands) around the underlying network topology and these technologies, to maximize information exchange. Finally, we illustrate this on a nine-city topology modeled on the DoubleZero network, a live programmable WAN equipped with both technologies, and show how the optimal schedules shift with the network's capabilities. Together, these can narrow the gap to the gold standard of colocated training.

分布式训练广域网网络优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。