arXiv:2503.03213stat.MLcs.LG2025-03被引 3

分析软门控混合专家模型的收敛速度,揭示专家结构对数据效率的影响。

Convergence Rates for Softmax Gating Mixture of Experts

  • 基于软门控机制,建立参数与专家估计的收敛性理论。
  • 强可辨识条件下,需多项式量级数据;线性专家则需指数级数据。
  • 为高效专家结构设计提供理论指导,适合模型优化研究者。

混合专家(MoE)通过将复杂任务动态分配给多个专用子模型(专家)来提升机器学习模型的效率与可扩展性。其核心是自适应的软门控机制,负责判断每个专家对输入的相关性并分配权重。尽管广泛应用,但关于软门控对MoE影响的系统性研究仍不足。本文针对标准软门控、稀疏化门控及分层软门控等变体,进行了参数估计与专家估计的收敛性分析。理论表明,在满足作者提出的 extit{强可辨识}条件(如常见的两层前馈网络)时,仅需多项式量级数据即可估计专家;而违反该条件的线性专家因内在参数交互(由偏微分方程描述),需指数级数据。所有结论均得到严格理论保证。

原文摘要 · Abstract (English)

Mixture of experts (MoE) has recently emerged as an effective framework to advance the efficiency and scalability of machine learning models by softly dividing complex tasks among multiple specialized sub-models termed experts. Central to the success of MoE is an adaptive softmax gating mechanism which takes responsibility for determining the relevance of each expert to a given input and then dynamically assigning experts their respective weights. Despite its widespread use in practice, a comprehensive study on the effects of the softmax gating on the MoE has been lacking in the literature. To bridge this gap in this paper, we perform a convergence analysis of parameter estimation and expert estimation under the MoE equipped with the standard softmax gating or its variants, including a dense-to-sparse gating and a hierarchical softmax gating, respectively. Furthermore, our theories also provide useful insights into the design of sample-efficient expert structures. In particular, we demonstrate that it requires polynomially many data points to estimate experts satisfying our proposed \emph{strong identifiability} condition, namely a commonly used two-layer feed-forward network. In stark contrast, estimating linear experts, which violate the strong identifiability condition, necessitates exponentially many data points as a result of intrinsic parameter interactions expressed in the language of partial differential equations. All the theoretical results are substantiated with a rigorous guarantee.

混合专家收敛分析模型效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。