arXiv:2502.05172cs.LGcs.AI2025-02ICML被引 30

MoE模型在内存限制下比密集模型更高效,可显著降低部署成本。

Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient

  • 基于激活参数、数据集大小等构建联合缩放定律
  • 280次实验验证:2.7亿活跃参数下内存效率优于密集模型
  • 为大规模训练提供可落地的MoE配置优化方案

Mixture of Experts(MoE)架构在大规模机器学习模型的研究与实际应用中显著提升了计算效率。然而,其在内存约束下的可扩展性与效率仍相对未被充分探索。本文提出稠密模型与MoE模型的联合缩放定律,纳入活跃参数数量、数据集规模和专家数量等关键因素。研究结果为在固定内存与算力预算下选择最优MoE配置提供了理论依据。令人意外的是,我们发现MoE模型在特定条件下可比稠密模型更具内存效率,颠覆了传统认知。为推导并验证缩放定律的理论预测,我们进行了超过280次实验,涉及最多27亿活跃参数及最多50亿总参数。这些结果为实际大规模训练场景中MoE模型的设计与部署提供了可操作的洞见。

原文摘要 · Abstract (English)

Mixture of Experts (MoE) architectures have significantly increased computational efficiency in both research and real-world applications of large-scale machine learning models. However, their scalability and efficiency under memory constraints remain relatively underexplored. In this work, we present joint scaling laws for dense and MoE models, incorporating key factors such as the number of active parameters, dataset size, and the number of experts. Our findings provide a principled framework for selecting the optimal MoE configuration under fixed memory and compute budgets. Surprisingly, we show that MoE models can be more memory-efficient than dense models, contradicting conventional wisdom. To derive and validate the theoretical predictions of our scaling laws, we conduct over 280 experiments with up to 2.7B active parameters and up to 5B total parameters. These results offer actionable insights for designing and deploying MoE models in practical large-scale training scenarios.

MoE内存效率缩放定律大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。