用开源框架提升早期AI硬件训练的可靠性与效率。
SIGMA: An AI-Empowered Training Stack on Early-Life Hardware
- 设计LTP系统应对早期硬件的不稳定与复杂并行问题。
- 2048块加速器训练200B模型,75天仅1次故障,MFU达21.08%。
- 适合关注低成本高可靠大模型训练的研究者与工程师。
越来越多的AI加速器被用于大规模训练,但早期生命周期的硬件面临三大挑战:频繁系统中断与未定义故障模式影响可靠性;数值误差与训练不稳定性威胁正确性与收敛性;并行优化复杂性与局部噪声不确定性降低效率。为此,SIGMA是面向早期AI硬件的大规模分布式训练开源栈,核心为专为早期硬件集群优化的LUCIA TRAINING PLATFORM(LTP)。自2025年3月上线以来,LTP显著提升训练可靠性与运维效率,过去五个月实现94.45%的有效加速器利用率,大幅减少节点回收与任务恢复时间。基于LTP,LUCIA TRAINING FRAMEWORK(LTF)成功使用2,048块加速器训练出SIGMA-MOE(200B MoE模型),达成21.08% MFU、顶尖下游准确率,且在75天内仅发生一次稳定性事件。这些成果不仅解决大规模训练的关键挑战,更树立了AI基础设施新标杆,为现有加速器栈提供稳健、经济的替代方案,显著推动AI能力与可扩展性。源码见https://github.com/microsoft/LuciaTrainingPlatform。
原文摘要 · Abstract (English)
An increasing variety of AI accelerators is being considered for large-scale training. However, enabling large-scale training on early-life AI accelerators faces three core challenges: frequent system disruptions and undefined failure modes that undermine reliability; numerical errors and training instabilities that threaten correctness and convergence; and the complexity of parallelism optimization combined with unpredictable local noise that degrades efficiency. To address these challenges, SIGMA is an open-source training stack designed to improve the reliability, stability, and efficiency of large-scale distributed training on early-life AI hardware. The core of this initiative is the LUCIA TRAINING PLATFORM (LTP), the system optimized for clusters with early-life AI accelerators. Since its launch in March 2025, LTP has significantly enhanced training reliability and operational productivity. Over the past five months, it has achieved an impressive 94.45% effective cluster accelerator utilization, while also substantially reducing node recycling and job-recovery times. Building on the foundation of LTP, the LUCIA TRAINING FRAMEWORK (LTF) successfully trained SIGMA-MOE, a 200B MoE model, using 2,048 AI accelerators. This effort delivered remarkable stability and efficiency outcomes, achieving 21.08% MFU, state-of-the-art downstream accuracy, and encountering only one stability incident over a 75-day period. Together, these advances establish SIGMA, which not only tackles the critical challenges of large-scale training but also establishes a new benchmark for AI infrastructure and platform innovation, offering a robust, cost-effective alternative to prevailing established accelerator stacks and significantly advancing AI capabilities and scalability. The source code of SIGMA is available at https://github.com/microsoft/LuciaTrainingPlatform.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。