arXiv:2502.03885cs.NIcs.DC2025-02被引 22

用光路交换收发器构建可扩展的超宽带数据中心网络,解决大模型训练通信瓶颈。

InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers

论文配图:InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers
图 1 · 摘自论文原文
  • 在收发器级集成光电路交换,实现动态点对多点通信
  • 成本降至NVL-72的31%,故障节点下GPU浪费率降超10倍
  • 适合超大规模AI训练集群,尤其关注成本与容错的团队

大模型训练依赖多维并行,高带宽域(HBD)对通信密集型并行(如张量并行)至关重要。现有架构存在可扩展性差、成本高、容错性弱等问题:以交换机为中心的方案(如NVL-72)扩展成本过高,以GPU为中心的方案(如TPUv3/Dojo)故障传播严重,混合方案(如TPUv4)仍存在较大故障扩散范围。本文提出InfiniteHBD,一种以收发器为中心的HBD架构,通过在每个收发器内嵌硅光子基光电路交换(OCS)实现连接与动态切换一体化,支持可重构的k跳环形拓扑与点对多点通信。该设计实现了数据中心级可扩展性、节点级故障隔离和健康GPU全带宽利用。关键创新包括光电路交换收发器(OCSTrx)、可重构环形拓扑及HBD-DCN编排算法。评估显示,InfiniteHBD将成本降至NVL-72的31%,GPU浪费率近零(比NVL-72和TPUv4低10倍以上),在7%节点故障率下跨机架流量接近零,模型每秒浮点运算利用率较NVIDIA DGX(8 GPU/节点)提升3.37倍。

原文摘要 · Abstract (English)

Scaling Large Language Model (LLM) training relies on multi-dimensional parallelism, where High-Bandwidth Domains (HBDs) are critical for communication-intensive parallelism like Tensor Parallelism. However, existing HBD architectures face fundamental limitations in scalability, cost, and fault resiliency: switch-centric HBDs (e.g., NVL-72) incur prohibitive scaling costs, while GPU-centric HBDs (e.g., TPUv3/Dojo) suffer from severe fault propagation. Switch-GPU hybrid HBDs (e.g., TPUv4) take a middle-ground approach, but the fault explosion radius remains large. We propose InfiniteHBD, a transceiver-centric HBD architecture that integrates connectivity and dynamic switching at the transceiver level by embedding Optical Circuit Switching (OCS) within each transceiver. It enables reconfigurable point-to-multipoint communication and scalable variable-size ring topologies. InfiniteHBD achieves datacenter-scale scalability without cost explosion, fault isolation at the node level, and full bandwidth utilization for healthy GPUs. Key innovations include a Silicon Photonic-based OCS transceiver (OCSTrx), a reconfigurable k-hop ring topology, and an HBD-DCN orchestration algorithm. The evaluation demonstrates that InfiniteHBD reduces cost to 31% of NVL-72, achieves a near-zero GPU waste ratio (over 10x lower than NVL-72 and TPUv4), maintains near-zero cross-ToR traffic under 7% node fault ratio, and improves Model FLOPs Utilization by 3.37x compared to NVIDIA DGX (8 GPUs/node).

大模型训练光通信数据中心网络容错架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。