arXiv:2410.02935stat.MLcs.LG2024-10被引 15

用拉普拉斯门控替代Softmax,提升分层专家模型的性能与专家专精度。

On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions

  • 采用拉普拉斯门控函数替代传统Softmax,解决参数耦合问题。
  • 在图像分类与多模态任务中,模型收敛更快且专家分工更明确。
  • 适合需要高效专家分工的大规模模型研发人员参考。

随着混合专家(MoE)架构在构建大规模基础模型中的日益重要,本文研究了分层混合专家(HMoE)这一特殊变体,其在处理复杂输入和提升特定任务性能方面表现优异。分析表明,相较于传统的Softmax门控,使用拉普拉斯门控函数在HMoE的两个层级均能有效消除由Softmax引入的不良参数交互,从而加速专家收敛并增强专家专一性。理论证明显示,该方法可显著改善模型训练动态。在多个场景下的实证验证支持上述结论,包括大规模多模态任务、图像分类以及潜在领域发现与预测任务,改进后的HMoE模型相较传统版本均有显著性能提升。

原文摘要 · Abstract (English)

With the growing prominence of the Mixture of Experts (MoE) architecture in developing large-scale foundation models, we investigate the Hierarchical Mixture of Experts (HMoE), a specialized variant of MoE that excels in handling complex inputs and improving performance on targeted tasks. Our analysis highlights the advantages of using the Laplace gating function over the traditional Softmax gating within the HMoE frameworks. We theoretically demonstrate that applying the Laplace gating function at both levels of the HMoE model helps eliminate undesirable parameter interactions caused by the Softmax gating and, therefore, accelerates the expert convergence as well as enhances the expert specialization. Empirical validation across diverse scenarios supports these theoretical claims. This includes large-scale multimodal tasks, image classification, and latent domain discovery and prediction tasks, where our modified HMoE models show great performance improvements compared to the conventional HMoE models.

专家模型门控机制分层结构模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。