分析百万级训练任务数据,揭示大模型集群可靠性瓶颈与优化方向。
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
- 基于11个月、400万任务数据,量化分析不同规模作业的故障率差异。
- 发现大任务最易失败,但小任务占多数,需纳入优化考量。
- 提出故障分类体系与有效训练时间比模型,助力系统级可靠性改进。
可靠性是大规模机器学习基础设施的核心挑战,尤其在模型与训练集群持续扩大的背景下。本文通过对两个大型多租户ML集群的运营分析,提供了量化评估、实践经验与对可靠性问题的深入见解。分析显示,尽管大任务最易受故障影响,但小任务占集群任务总量的多数,应纳入优化目标。我们识别关键工作负载特性,跨集群比较并揭示了推动大规模机器训练所需的关键可靠性要求。引入故障分类体系与核心可靠性指标,分析来自两个前沿ML环境的11个月数据,涵盖超过400万次任务和1.5亿+ A100 GPU小时。基于数据构建故障模型,预测不同GPU规模下的平均失效时间。进一步提出一种估算有效训练时间比的方法,作为任务参数的函数,并用于评估软件缓解措施在大规模下的有效性。本工作为提升人工智能超算集群可靠性提供重要洞见与未来研究方向,强调需构建灵活、无工作负载依赖且具备可靠性感知能力的基础设施、系统软件与算法。
原文摘要 · Abstract (English)
Reliability is a fundamental challenge in operating large-scale machine learning (ML) infrastructures, particularly as the scale of ML models and training clusters continues to grow. Despite decades of research on infrastructure failures, the impact of job failures across different scales remains unclear. This paper presents a view of managing two large, multi-tenant ML clusters, providing quantitative analysis, operational experience, and our own perspective in understanding and addressing reliability concerns at scale. Our analysis reveals that while large jobs are most vulnerable to failures, smaller jobs make up the majority of jobs in the clusters and should be incorporated into optimization objectives. We identify key workload properties, compare them across clusters, and demonstrate essential reliability requirements for pushing the boundaries of ML training at scale. We hereby introduce a taxonomy of failures and key reliability metrics, analyze 11 months of data from two state-of-the-art ML environments with 4 million jobs and over 150 million A100 GPU hours. Building on our data, we fit a failure model to project Mean Time to Failure for various GPU scales. We further propose a method to estimate a related metric, Effective Training Time Ratio, as a function of job parameters, and we use this model to gauge the efficacy of potential software mitigations at scale. Our work provides valuable insights and future research directions for improving the reliability of AI supercomputer clusters, emphasizing the need for flexible, workload-agnostic, and reliability-aware infrastructure, system software, and algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。