arXiv:2605.21187cs.NIcs.AI2026-05

为超大规模AI训练设计高速网络,实现低延迟高稳定。

High-speed Networking for Giga-Scale AI Factories

论文配图:High-speed Networking for Giga-Scale AI Factories
图 1 · 摘自论文原文
  • 用拓扑并行替代层级结构,提升网络扩展性。
  • 实测达到理论带宽98%,延迟抖动极小。
  • 适合大模型训练等对网络响应要求极高的场景。

随着分布式模型训练扩展至数十万张GPU,横向扩展网络面临前所未有的性能与效率挑战。NVIDIA Spectrum-X以全新设计实现可预测、稳定的网络表现,具备高利用率和低延迟。本文提出Spectrum-X多平面架构,以拓扑并行取代传统层级深度,并在网卡与交换机中引入硬件加速负载均衡,实现微秒级对高度动态网络条件的快速响应,满足AI训练需求。我们阐述了设计动机、原则、评估方法及在先进基准测试中的表现,并总结了在大规模系统部署与调试中获得的经验。评估显示,该网络在三个核心维度表现卓越:理论线速率的98%且无抖动延迟;支持多租户并发工作负载的强隔离;带宽按容量比例增长,10%链路故障仅导致7%延迟增加;在大语言模型训练中能快速应对主机和链路故障。

原文摘要 · Abstract (English)

As distributed model training scales to span hundreds of thousands of GPUs, scale-out networks face unprecedented performance and efficiency demands. NVIDIA Spectrum-X Ethernet has been designed from the ground up to achieve predictable and stable network performance with high utilization and low latency. This paper presents the Spectrum-X multiplane architecture, which replaces hierarchical depth with topological parallelism, and introduces hardware-accelerated load balancing in NICs and switches as the key architectural approach to provide fast reaction to highly dynamic network conditions at the microsecond timescales that AI training workloads demand. We describe the motivation, design principles, evaluation methodology and performance on state-of-the-art benchmarks, as well as the lessons we learned from deploying and debugging Spectrum-X networks in large-scale systems. Our evaluation highlights production-grade AI infrastructure performance across three core dimensions: 98% of the theoretical line rate with low jitter-free latency; strong cross-tenant isolation for concurrent workloads; robust, capacity-proportional bisection bandwidth and 7% latency increase for 10% fabric link failures; and rapid reaction to host and fabric link flaps during LLM training workloads.

AI基础设施高速网络分布式训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。