arXiv:2505.10860cs.LGstat.ML2025-05被引 3

解析DeepSeekMoE的共享专家与归一化门控机制为何提升模型效率

On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating

  • 从统计角度分析共享专家和归一化门控的理论优势
  • 实验证明该设计显著提升样本效率,降低训练波动
  • 适合关注大模型架构设计与门控机制优化的研究者

混合专家(MoE)方法是当前大型语言模型架构的核心组件,包括近期的DeepSeek系列模型。相比其他MoE实现,DeepSeekMoE的独特之处在于采用了共享专家策略和归一化sigmoid门控机制。尽管该结构在DeepSeek系列模型的成功中起关键作用,但对其共享专家策略的理论依据研究有限,且归一化sigmoid门控尚未被系统探讨。为此,本文从统计视角对DeepSeekMoE的这两项特征进行系统性理论研究。通过专家估计任务的收敛性分析,揭示了共享专家策略和归一化sigmoid门控在样本效率上的增益,为专家与门控结构的设计提供了重要洞见。为验证理论发现,我们在合成数据和真实世界数据集(视觉-语言建模任务)上进行了多项实验。最后,我们对路由器行为进行了广泛实证分析,涵盖路由器饱和度、变化率及专家利用率等指标。

原文摘要 · Abstract (English)

Mixture of experts (MoE) methods are a key component in most large language model architectures, including the recent series of DeepSeek models. Compared to other MoE implementations, DeepSeekMoE stands out because of two unique features: the deployment of a shared expert strategy and of the normalized sigmoid gating mechanism. Despite the prominent role of DeepSeekMoE in the success of the DeepSeek series of models, there have been only a few attempts to justify theoretically the value of the shared expert strategy, while its normalized sigmoid gating has remained unexplored. To bridge this gap, we undertake a comprehensive theoretical study of these two features of DeepSeekMoE from a statistical perspective. We perform a convergence analysis of the expert estimation task to highlight the gains in sample efficiency for both the shared expert strategy and the normalized sigmoid gating, offering useful insights into the design of expert and gating structures. To verify empirically our theoretical findings, we carry out several experiments on both synthetic data and real-world datasets for (vision) language modeling tasks. Finally, we conduct an extensive empirical analysis of the router behaviors, ranging from router saturation, router change rate, to expert utilization.

MoE门控机制大模型统计分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。