发现专家混合模型可通过线性路径连接,支持高效集成与泛化。
On Linear Mode Connectivity of Mixture-of-Experts Architectures
- 提出匹配算法对齐独立训练的专家模型,实现参数空间线性连通。
- 在密集、稀疏及共享专家等多种配置中均验证线性模式连通性。
- 为理解深度模型优化动态和损失几何提供新视角,适合研究者参考。
线性模式连通性(Linear Mode Connectivity, LMC)是神经网络损失曲面中的显著现象:独立训练的模型可通过参数空间中的线性路径连接,且路径上损失始终较低,这挑战了非凸优化的传统认知,并对模型集成、泛化能力及神经网络损失几何的理解具有重要意义。受标准神经网络中LMC研究的启发,我们系统研究了专家混合(Mixture-of-Experts, MoE)架构中的该现象。MoE通过可学习的门控机制组合多个专家网络,具有良好的可扩展性和计算效率。我们首先分析了密集与稀疏门控两种情形,证明MoE固有的对称性由专家组件与门控函数上的置换共同决定。基于此,我们提出一种匹配算法,用于对齐独立训练的MoE模型,从而发现其线性连通路径。最后,我们在多种MoE配置(包括密集、稀疏、共享专家)下,跨不同模型设置与多尺度、多模态数据集上,通过实验验证了LMC的存在。结果表明,即使在复杂结构中,独立训练的MoE仍可通过低损失路径连接,揭示了深层模型功能景观与优化动力学的本质特性。
原文摘要 · Abstract (English)
Linear Mode Connectivity (LMC) is a notable phenomenon in the loss landscapes of neural networks, wherein independently trained models have been observed to be connected--up to permutation symmetries--by linear paths in parameter space along which the loss remains consistently low. This observation challenges classical views of non-convex optimization and has implications for model ensembling, generalization, and our understanding of neural loss geometry. Inspired by recent studies on LMC in standard neural networks, we systematically investigate this phenomenon within Mixture-of-Experts (MoE) architectures--a class of models known for their scalability and computational efficiency, which combine traditional neural networks--referred to as experts--through a learnable gating mechanism. We begin by conducting a comprehensive analysis of both dense and sparse gating regimes, demonstrating that the symmetries inherent to MoE architectures are fully characterized by permutations acting on both the expert components and the gating function. Building on these foundational findings, we propose a matching algorithm that enables alignment between independently trained MoEs, thereby facilitating the discovery of LMC. Finally, we empirically validate the presence of LMC using our proposed algorithm across diverse MoE configurations--including dense, sparse, and shared-expert variants--under a wide range of model settings and datasets of varying scales and modalities. Our results confirm the existence of LMC in MoE architectures and offer fundamental insights into the functional landscape and optimization dynamics of deep learning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。