arXiv:2607.28308cs.LG2026-07

发现稀疏专家模型中专家重叠反而提升性能,挑战传统几何互补假设。

Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing

  • 提出新指标分离路由中的几何重叠、候选质量与上下文交互。
  • 实验证明专家子空间高度重叠,但实际路由仍优于随机匹配。
  • 即使专家相似,多专家协作仍有效,适合研究模型路由机制者。

稀疏专家混合(MoE)语言模型将每个标记分配给多个专家,传统观点认为其优势源于专家贡献不同表征方向。现有证据常混淆路由一致性、候选质量与候选-上下文交互。本文通过专家子空间分离指数(ESSI)、匹配路由残差及前缀控制的2×2因子实验,区分这些量。冻结路由干预和受控Top-k研究评估功能价值。三个对比揭示:第一,在六种MoE架构中,专家子空间显著重叠,但实际路由比匹配替代方案更优;第二,在OLMoE、Mixtral和DeepSeek的39个因子单元中,所选候选解释残差表示比最强未选竞争者更佳,但实际前缀使该优势缩小——所有交互均为负,95%置信区间均低于零;第三,这种几何缩小不意味功能冗余:在39次冻结路由比较中,24次添加后期专家提升下一个词预测性能,其余15次结果不确定;受控训练研究也显示三种种子下Top-2优于Top-1。我们称此现象为‘相干重叠’:路由从共享几何邻域中选择相关专家,而无需线性不重叠即可实现有效多专家计算。分离这些量澄清了为何仅凭几何相似性无法判断冗余或剪枝价值。

原文摘要 · Abstract (English)

Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction. We distinguish these quantities using an Expert Subspace Separation Index (ESSI), matched-route residuals, and a prefix-controlled $2\times2$ factorial; frozen-route interventions and a controlled Top-$k$ study assess functional value. Three paired contrasts organize the findings. First, across six MoE architectures, expert subspaces overlap substantially, yet actual routes explain token representations better than matched alternatives. Second, across the 39 factorial cells in OLMoE, Mixtral, and DeepSeek, the selected candidate explains more of the residual representation than the strongest unselected rival in every cell, yet the actual prefix narrows this advantage throughout: all interactions are negative, and every 95% confidence interval lies below zero. Third, this geometric narrowing does not imply functional redundancy: adding later experts improves next-token prediction in 24 of 39 frozen-route comparisons, while the other 15 estimates are inconclusive; a controlled training study also favors Top-2 over Top-1 in all three seeds. We call this joint pattern coherent overlap: routing selects token-relevant experts from a shared geometric neighborhood, while useful multi-expert computation persists without disjoint linear coverage. Separating these quantities clarifies why geometric similarity alone cannot determine redundancy or pruning value.

MoE专家模型路由机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。