arXiv:2604.09175cs.LGcs.AI2026-04

揭示MoE Transformer的泛化与扩展规律,明确活跃参数的作用

Generalization and Scaling Laws for Mixture-of-Experts Transformers

  • 分离输入活跃容量与路由复杂性,建立基于覆盖数的泛化界
  • 证明误差可随活跃参数或专家数量增加而降低,取决于瓶颈所在
  • 给出模型规模、数据量与算力最优权衡的理论依据,适合研究者参考

我们建立了Mixture-of-Experts(MoE)Transformer的泛化与扩展理论,清晰区分了每输入的活跃容量与路由组合复杂性。通过固定路由模式并进行并集边界估计,推导出一种超范数覆盖数界,其度量熵依赖于活跃参数预算,并包含特定于MoE的路由开销。结合平方损失的标准经验风险最小化分析,在$d$维流形数据模型和$C^β$目标下,表明当正确考虑活跃参数后,逼近与估计之间的权衡与密集网络一致。进一步证明了MoE架构的构造性逼近定理:在构造条件下,误差可通过增加活跃容量或专家数量来降低,取决于主导瓶颈。由此推导出模型规模、数据规模及算力最优权衡的神经扩展规律。总体而言,我们的结果为理解MoE扩展提供了透明的统计基准,澄清了哪些行为是理论保证的,哪些需依赖数据相关的路由结构或优化动力学。

原文摘要 · Abstract (English)

We develop a theory of generalization and scaling for Mixture-of-Experts (MoE) Transformers that cleanly separates \emph{active} per-input capacity from routing combinatorics. By conditioning on fixed routing patterns and union-bounding across them, we derive a sup-norm covering-number bound whose metric entropy scales with the active parameter budget and incurs a MoE-specific routing overhead. Combined with a standard ERM analysis for squared loss, this yields a generalization bound under a $d$-dimensional manifold data model and $C^β$ targets, showing that approximation and estimation trade off as in dense networks once active parameters are accounted for appropriately. We further prove a constructive approximation theorem for MoE architectures, showing that, under the approximation construction, error can decrease either by scaling active capacity or by increasing the number of experts, depending on the dominant bottleneck. From these results we derive neural scaling laws for model size, data size, and compute-optimal tradeoffs. Overall, our results provide a transparent statistical reference point for reasoning about MoE scaling, clarifying which behaviors are certified by worst-case theory and which must arise from data-dependent routing structure or optimization dynamics.

MoE扩展规律泛化理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。