arXiv:2409.13198cs.CLcs.LG2024-09被引 5

局部SGD在大模型训练中表现不逊于传统方法,适合分布式场景。

Exploring Scaling Laws for Local SGD in Large Language Model Training

  • 采用局部SGD实现松散连接设备上的分布式训练
  • 同等资源下性能接近单一大集群训练
  • 适用于多集群和边缘计算等实际场景

本文研究了在大语言模型训练中局部SGD的扩展规律,这是一种可在松散连接设备上进行训练的分布式优化算法。通过大量实验,我们发现,在模型参数、数据集和计算资源相当的情况下,局部SGD能达到与传统方法相当的性能。此外,我们探讨了局部SGD在多集群设置和边缘计算环境中的应用。研究揭示了有效进行多集群大模型训练的必要条件,并分析了利用边缘计算资源在大模型训练中的潜力与局限。结果表明,局部SGD可作为单一大集群训练的可行替代方案。

原文摘要 · Abstract (English)

This paper investigates scaling laws for local SGD in LLM training, a distributed optimization algorithm that facilitates training on loosely connected devices. Through extensive experiments, we show that local SGD achieves competitive results compared to conventional methods, given equivalent model parameters, datasets, and computational resources. Furthermore, we explore the application of local SGD in various practical scenarios, including multi-cluster setups and edge computing environments. Our findings elucidate the necessary conditions for effective multi-cluster LLM training and examine the potential and limitations of leveraging edge computing resources in the LLM training process. This demonstrates its viability as an alternative to single large-cluster training.

分布式训练大模型局部SGD边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。