arXiv:2510.07205cs.LGmath.OC2025-10被引 1

首次证明软路由专家模型可收敛,专家指导路由学习。

Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts

  • 在学生-教师框架下,用非线性路由和专家联合训练。
  • 适度过参数化时,路由学习被专家引导,可恢复教师参数。
  • 剪枝后微调能全局最优,适合研究模型优化机制者。

Mixture-of-Experts(MoE)架构已成为现代AI系统的核心。其动态路由将输入分配给专用专家,输出通过加权求和聚合。尽管广泛应用,现有理论仅限于独立优化专家与路由,或仅限于精心构造数据集的Top-1路由场景。本文首次在学生-教师框架下,为非线性路由与专家的软路由MoE模型提供联合训练的收敛性保证。证明在适度过参数化条件下,学生网络经历特征学习阶段,路由器的学习过程受专家引导,逐步恢复教师参数。此外,训练后剪枝可有效移除冗余神经元,并通过可证明收敛的微调达到全局最优。本分析为理解MoE优化景观提供了新视角。

原文摘要 · Abstract (English)

Mixture-of-Experts (MoE) architectures have emerged as a cornerstone of modern AI systems. In particular, MoEs route inputs dynamically to specialized experts whose outputs are aggregated through weighted summation. Despite their widespread application, theoretical understanding of MoE training dynamics remains limited to either separate expert-router optimization or only top-1 routing scenarios with carefully constructed datasets. This paper advances MoE theory by providing convergence guarantees for joint training of soft-routed MoE models with non-linear routers and experts in a student-teacher framework. We prove that, with moderate over-parameterization, the student network undergoes a feature learning phase, where the router's learning process is ``guided'' by the experts, that recovers the teacher's parameters. Moreover, we show that a post-training pruning can effectively eliminate redundant neurons, followed by a provably convergent fine-tuning process that reaches global optimality. To our knowledge, our analysis is the first to bring novel insights in understanding the optimization landscape of the MoE architecture.

MoE优化理论专家系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。