用机器学习自动优化超大规模模型训练速度
CubicML: Automated ML for Large ML Systems Co-design with ML Prediction of Performance
- 用机器学习模型预测系统性能,提升调优效率
- 在730亿至4050亿参数模型上显著加速训练
- 适合需要高效训练超大模型的工业级团队
扩大深度学习模型规模已被证明能有效提升机器学习模型的智能水平,尤其在工业推荐系统和大语言模型中表现显著。大规模分布式机器学习系统与算法的协同设计(以最大化训练性能)是其成功的关键。随着模型规模扩大,协同设计的超参数数量迅速增加,导致难以高效找到最优配置。本文提出CubicML,利用机器学习模型作为代理,预测训练性能以实现搜索效率与性能建模灵活性的平衡。实验表明,CubicML可在Meta内部的730亿参数广告推荐模型和高达4050亿参数的大语言模型上有效优化训练速度。
原文摘要 · Abstract (English)
Scaling up deep learning models has been proven effective to improve intelligence of machine learning (ML) models, especially for industry recommendation models and large language models. The co-design of large distributed ML systems and algorithms (to maximize training performance) plays a pivotal role for its success. As it scales, the number of co-design hyper-parameters grows rapidly which brings challenges to feasibly find the optimal setup for system performance maximization. In this paper, we propose CubicML which uses ML to automatically optimize training performance of large distributed ML systems. In CubicML, we use an ML model as a proxy to predict the training performance for search efficiency and performance modeling flexibility. We proved that CubicML can effectively optimize training speed of in-house ads recommendation models with 73 billion parameters and large language models up to 405 billion parameters at Meta.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。