数据移动瓶颈将限制大模型训练规模,三年内可能触及根本极限。
Data movement limits to frontier model training
- 建立分布式训练理论模型,分析稠密与稀疏训练的可扩展性。
- 超10^28 FLOP训练运行将因数据移动导致硬件利用率下降。
- 适合关注大模型训练瓶颈与算力极限的研究者阅读。
我们提出一种分布式训练的理论模型,并用于分析稠密与稀疏训练的可扩展性。在基准假设下,若训练时长为三个月,当训练量超过约10^28 FLOP时,数据移动瓶颈将显著降低硬件利用率,这一规模比当前最大训练量高出两个数量级,表明在现有增长速率下,三年内将面临根本性扩展障碍。即使在低利用率情况下,超过10^31 FLOP的训练也难以实现。然而,更激进的批量大小扩展和/或更短更宽的模型结构(如更宽的网络形状)若能实现,则有望支持更大规模的训练。
原文摘要 · Abstract (English)
We present a theoretical model of distributed training, and use it to analyze how far dense and sparse training runs can be scaled. Under our baseline assumptions, given a three month training duration, data movement bottlenecks begin to significantly lower hardware utilization for training runs exceeding about $10^{28}$ FLOP, two orders of magnitude above the largest training run to date, suggesting the arrival of fundamental barriers to scaling in three years given recent rates of growth. A training run exceeding about $10^{31}$ FLOP is infeasible even at low utilization. However, more aggressive batch size scaling and/or shorter and fatter model shapes, if achievable, have the potential to permit much larger training runs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。