arXiv:2410.05661cs.LGcs.AI2024-10EMNLP被引 24

发现MoE模型与密集模型遵循相似的缩放规律,且在相同算力下表现更优。

Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models

  • 对比密集模型与MoE模型的缩放规律,验证幂律框架适用性
  • 相同训练算力下,MoE模型测试损失更低,泛化能力更强
  • 为MoE模型训练部署提供可迁移的优化策略参考

大规模语言模型(LLMs)的缩放是提升训练与部署效率的关键研究方向。本文通过理论分析与大量实验,包括一致的损失缩放、最优批量大小与学习率缩放,以及资源分配策略缩放,探究了密集模型与专家混合(MoE)模型之间缩放规律的可转移性与差异性。结果表明,幂律缩放框架同样适用于MoE模型,说明尽管架构不同,其缩放行为的基本原理仍保持一致。此外,MoE模型展现出更优的泛化能力,在相同训练计算预算下测试损失更低。这些发现揭示了MoE模型在缩放一致性与泛化能力上的优势,为优化其训练与部署策略提供了新视角。

原文摘要 · Abstract (English)

The scaling of large language models (LLMs) is a critical research area for the efficiency and effectiveness of model training and deployment. Our work investigates the transferability and discrepancies of scaling laws between Dense Models and Mixture of Experts (MoE) models. Through a combination of theoretical analysis and extensive experiments, including consistent loss scaling, optimal batch size and learning rate scaling, and resource allocation strategies scaling, our findings reveal that the power-law scaling framework also applies to MoE Models, indicating that the fundamental principles governing the scaling behavior of these models are preserved, even though the architecture differs. Additionally, MoE Models demonstrate superior generalization, resulting in lower testing losses with the same training compute budget compared to Dense Models. These findings indicate the scaling consistency and transfer generalization capabilities of MoE Models, providing new insights for optimizing MoE Model training and deployment strategies.

大模型缩放定律MoE泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。