大模型用细粒度专家结构,能更高效地训练和提升性能。
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
- 采用更多小专家的细粒度专家混合架构
- 560亿参数模型验证损失更低,下游任务准确率更高
- 为超大规模模型设计提供实证指导
混合专家(MoE)架构已成为高效扩展大语言模型的关键。细粒度MoE方法——即使用更多、更小的专家——在提升模型收敛速度与质量方面展现出潜力。本文提出一系列训练方案,并对细粒度MoE进行全面实证评估,直接对比其与标准MoE配置在总参数达560亿(活跃参数170亿)规模下的扩展特性。研究涵盖收敛速度、下游基准表现及实际训练考量等多个方面。结果表明,在最大规模下,细粒度MoE实现了更低的验证损失和更高的下游任务准确率。本研究为未来大规模模型开发中应用细粒度MoE提供了实证基础与实用建议。
原文摘要 · Abstract (English)
Mixture of Experts (MoE) architectures have emerged as pivotal for scaling Large Language Models (LLMs) efficiently. Fine-grained MoE approaches - utilizing more numerous, smaller experts - have demonstrated potential in improving model convergence and quality. This work proposes a set of training recipes and provides a comprehensive empirical evaluation of fine-grained MoE, directly comparing its scaling properties against standard MoE configurations for models with up to 56B total (17B active) parameters. We investigate convergence speed, model performance on downstream benchmarks, and practical training considerations across various setups. Overall, at the largest scale we show that fine-grained MoE achieves better validation loss and higher accuracy across a set of downstream benchmarks. This study offers empirical grounding and practical insights for leveraging fine-grained MoE in the development of future large-scale models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。