arXiv:2508.06617cs.LGcs.AI2025-08

统一稀疏与密集大模型的缩放规律,助力资源最优配置

Generalizing Scaling Laws for Dense and Sparse Large Language Models

  • 提出适用于稠密和稀疏模型的通用缩放定律
  • 在同等算力下,对DeepSeek-V3等超大规模MoE模型更精准预测性能
  • 可反推最优模型规模或最佳稀疏度,适配研发与部署场景

尽管大型语言模型(LLM)取得进展,但预训练时最优模型规模的预测或资源分配仍具挑战。现有缩放定律多针对特定架构(稠密或稀疏),本文重新审视已有经验缩放定律,提出一种通用缩放定律,构建适用于稠密与稀疏大模型的统一框架。我们评估并对比了所提定律与现有定律,证明其能准确捕捉已有缩放行为。进一步通过等FLOP对比,验证了该定律在Mixture-of-Experts(MoE)类超大规模模型如DeepSeek-V3上的有效性。所提定律可用于根据给定稀疏度估计最优模型超参数(模型规模、训练数据量、计算量),或根据给定超参数识别最优稀疏度。

原文摘要 · Abstract (English)

Despite recent advancements of large language models (LLMs), optimally predicting the model size for LLM pretraining or allocating optimal resources still remains a challenge. Several efforts have addressed the challenge by proposing different empirical scaling laws, but almost all of them are architecture-specific (dense or sparse). In this work we revisit existing empirical scaling laws and propose a generalized scaling law to provide a unified framework that is applicable to both dense and sparse large language models. We evaluate and compare our proposed scaling law with existing scaling laws and demonstrate that our proposed scaling law captures the scaling behavior of existing scaling laws. Further, we show an IsoFLOP comparison between our proposed scaling law and the state-of-the-art scaling law to illustrate the effectiveness of our proposed scaling law for Mixture-of-Expert (MoE)-based very large LLMs like DeepSeek-V3. Our proposed scaling law can be used to estimate the best model hyperparameters (Model size, Tokens and Compute) for a given sparsity or to identify the optimal sparsity for the given model hyperparameters.

大模型缩放定律MoE资源优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。