arXiv:2605.14200cs.LGstat.ML2026-05被引 1

提出新型参数化方法,让专家模型在扩大规模时更稳定、性能持续提升。

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization

论文配图:How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization
图 1 · 摘自论文原文
  • 基于动力学平均场理论,推导出适应不同缩放方式的最优参数初始化。
  • 新方法在多尺度下实现学习率迁移与性能单调增长,突破旧方法局限。
  • 适合研究大规模语言模型缩放规律的学者,尤其关注高效专家架构设计者。

当前前沿大语言模型广泛采用混合专家(MoE)架构。尽管已有实证进展,但对超参数如何随网络宽度 $N$、专家宽度 $N_e$、专家数量 $M$、稀疏度 $K$ 及深度 $L$ 进行合理缩放以保证稳定性和最佳性能,仍缺乏系统性理解。本文分析三种缩放范式:(I) $Nowtie N_e$,(II) $Nowtie Mowtie K$,(III) $N, N_e, M, K$ 全比例缩放。针对每种范式,构建了新的动力学平均场理论(DMFT)描述训练动态。在此框架下,推导出满足最大更新(μ)准则的SGD和Adam参数化。然而发现该μP方法无法保证规模扩展时性能单调提升或学习率迁移鲁棒性。问题根源在于聚合动态中的尺度依赖可观测量,从而提出更强的‘最大缩放稳定性’准则。基于此,我们为所有三种范式分别推导出最大化缩放稳定参数化(MSSP),并通过独立的DMFT分析刻画其独特极限动态。实验验证,MSSP在跨范式场景下均能实现学习率迁移鲁棒性和性能随规模单调提升。结合现有深度缩放理论,本工作给出了完整的MoE架构缩放方案,涵盖宽度、深度、专家宽度及专家数量。

原文摘要 · Abstract (English)

Recent frontier large language models predominantly rely on Mixture-of-Experts (MoE) architectures. Despite empirical progress, there is still no principled understanding of how hyperparameters should scale with network width $N$, expert width $N_e$, number of experts $M$, sparsity $K$, and depth $L$ to ensure both stability and optimal performance at scale. We take a principled step toward resolving this gap by analyzing three different scaling regimes: (I) co-scaling $N\asymp N_e$, (II) co-scaling $N\asymp M\asymp K$, and (III) full proportional scaling of $N, N_e, M$, and $K$. For each regime, we develop a novel Dynamical Mean Field Theory (DMFT) description of the limiting training dynamics of MoEs that provides a formal foundation for our analysis. Within this framework, we derive the unique parameterization for SGD and Adam satisfying all maximal-update ($μ$) desiderata. We then show that the resulting $μ$P prescription does not reliably induce monotonic improvement with scale or robust learning-rate transfer. We trace these pathologies to scale-dependent observables in the aggregation dynamics, which motivates a refined set of desiderata that we term maximal scale stability. Guided by this principle, we derive a Maximally Scale-Stable Parameterization (MSSP) for both SGD and Adam in all three scaling regimes, and characterize the corresponding limiting dynamics - qualitatively distinct from the $μ$P limit - through a separate DMFT analysis. Experiments verify that MSSP robustly recovers learning rate transfer and monotonic improvement with scale across regimes. Combined with existing depth-scaling theory, these results provide a complete scaling prescription for MoE architectures as a function of width, depth, expert width, and number of experts.

MoE参数化缩放深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。