arXiv:2505.23105cs.LGcs.NI2025-05被引 1

LUMION通过光网络快速替换故障GPU,1秒内恢复训练,提升2倍效率。

LUMION: Fast Fault Recovery for ML Jobs Using Programmable Optical Fabrics

  • 用可编程光互连动态接入备用GPU,避免整机迁移。
  • 故障后1秒内完成替换与重启,吞吐量提升近2倍。
  • 适合高可用性ML训练场景,尤其看重恢复速度的团队。

现代ML数据中心中,加速器故障时,运维通常将受影响的训练或推理任务迁移到全新机架,虽保持网络性能,但需预留完整机架空闲加速器,资源浪费严重。本文提出LUMION,一种新型可重构光互连架构,用于机架内加速器互联。不同于整体迁移,LUMION在故障发生时动态将备用加速器接入当前工作负载,维持性能且无需昂贵迁移。我们构建了端到端硬件原型,实验显示其在微调Llama 3.2时,可在约1秒内完成故障GPU替换与任务重启。替换后,相比传统电互连机架,LUMION实现更高跨GPU带宽,使微调吞吐量接近提升2倍。

原文摘要 · Abstract (English)

When accelerators fail in modern ML datacenters, operators migrate the affected ML training or inference jobs to entirely new racks. This approach, while preserving network performance, is highly inefficient, requiring datacenters to reserve full racks of idle accelerators for fault tolerance. In this paper, we address this resource inefficiency by introducing LUMION, a novel reconfigurable optical fabric for connecting accelerators within a datacenter rack. Instead of migrating entire ML jobs, LUMION dynamically integrates spare accelerators into ongoing workloads as failures occur, thereby maintaining consistent performance without costly migrations. We show the benefits of LUMION by building an end-to-end hardware prototype. Our experiments fine-tune Llama 3.2 and show that LUMION swaps a failed GPU with a healthy one and restarts the ML job within ~ 1 second of the failure. LUMION achieves higher inter-GPU bandwidth compared to traditional electrical racks after replacing failed accelerators with spare ones, leading to nearly 2X improvement in fine-tuning throughput.

故障恢复光互连GPU调度ML训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。