用可编程光网络提升多租户AI服务器的带宽与容错能力
Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML
- 通过可编程光互连重构服务器内加速器拓扑,动态优化资源分配
- 多租户场景下带宽提升66%,计算碎片减少70%,训练吞吐量提高1.72倍
- 故障芯片1.2秒内自动替换,适合高可用、大规模AI训练系统
我们提出Morphlux,一种用于服务器级加速器互联的可编程光网络。实验表明,在现有基于环形拓扑的ML数据中心中引入Morphlux,可使租户计算资源的带宽提升最高达66%,计算碎片减少最多70%,并显著缩小芯片故障的影响范围。我们构建了Morphlux的端到端硬件原型,验证其性能优势:在硬件测试平台中快速编程,可实现故障加速器芯片在1.2秒内被健康芯片替代,整体模型训练吞吐量提升1.72倍。
原文摘要 · Abstract (English)
We develop Morphlux, a server-scale programmable photonic fabric to interconnect accelerators within servers. We show that augmenting state-of-the-art torus-based ML data-centers with Morphlux can improve the bandwidth of tenant compute allocations by up to 66%, reduce compute fragmentation by up to 70%, and minimize the blast radius of chip failures. We develop a novel end-to-end hardware prototype of Morphlux to demonstrate these performance benefits which translate to 1.72X improvement in training throughput of ML models. By rapidly programming the server-scale fabric in our hardware testbed, Morphlux can replace a failed accelerator chip with a healthy one in 1.2 seconds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。