arXiv:2411.13055cs.LGcs.DC2024-11被引 17

大规模训练中,硬件扩展快到一定限度后效率反而下降。

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

  • 研究不同硬件配置与并行策略对大模型训练的影响
  • 超过特定规模后,通信开销使原被认为低效的策略反而更优
  • 即使优化配置,增加加速器数量也很快出现边际收益递减

近年来神经网络模型能力的显著提升依赖于模型规模、训练数据量和计算资源的持续扩大。为训练现代应用所需的超大规模模型(如大语言模型),训练过程需在数万块硬件加速器(如GPU)上分布式进行,涉及大规模计算集群中的计算与通信协调。本文通过大规模实证研究,分析了不同模型规模、硬件配置及分布式并行策略下大模型训练的性能表现。结果表明:(1)在特定规模以上,某些分布式通信策略带来的开销使得原本被认为次优的并行策略反而更优;(2)即使硬件与并行策略均优化,继续增加加速器数量仍会迅速产生收益递减,说明每单位电力或GPU小时的边际性能提升低下。

原文摘要 · Abstract (English)

Dramatic increases in the capabilities of neural network models in recent years are driven by scaling model size, training data, and corresponding computational resources. To develop the exceedingly large networks required in modern applications, such as large language models (LLMs), model training is distributed across tens of thousands of hardware accelerators (e.g. GPUs), requiring orchestration of computation and communication across large computing clusters. In this work, we demonstrate that careful consideration of hardware configuration and parallelization strategy is critical for effective (i.e. compute- and cost-efficient) scaling of model size, training data, and total computation. We conduct an extensive empirical study of the performance of large-scale LLM training workloads across model size, hardware configurations, and distributed parallelization strategies. We demonstrate that: (1) beyond certain scales, overhead incurred from certain distributed communication strategies leads parallelization strategies previously thought to be sub-optimal in fact become preferable; and (2) scaling the total number of accelerators for large model training quickly yields diminishing returns even when hardware and parallelization strategies are properly optimized, implying poor marginal performance per additional unit of power or GPU-hour.

大模型训练分布式硬件扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。