arXiv:2606.19036cs.LG2026-06被引 1

解析稀疏专家混合模型的不连续性,提出平滑机制提升稳定性与性能。

Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts

论文配图:Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts
图 1 · 摘自论文原文
  • 按专家切换时并列数量分类不连续性,发现低阶不连续集占主导体积。
  • 随机路径几乎必然在有限时间内击中一阶不连续面,且停留时间有界。
  • 设计轻量级平滑方法,增强模型连续性,实测提升语言与视觉任务表现。

稀疏专家混合(SMoE)架构广泛应用于先进语言与视觉模型中,通过条件路由实现大规模网络扩展。然而,其基于Top-k的专家选择导致映射本身存在固有不连续性:即使输入极其接近,也可能激活完全不同专家组合,产生显著差异输出。本文从几何与随机角度严格分析此类不连续性,首先按切换事件中并列专家数对不连续性进行分级;利用测度论切片论证,建立厚化不连续面的渐近体积估计,表明低阶不连续集主导空间,高阶者相对体积可忽略。进一步建模输入空间中的随机扰动为扩散过程,证明路径必然在有限时间内遭遇不连续,且首次击中几乎必然发生在一阶不连续面,给出明确的时间概率上界。还推导了路径在各类不连续邻域的驻留时间上限。理论表明输入更可能靠近低阶不连续区域。受此启发,提出一种可直接应用于现有SMoE的简单平滑机制,软性引入邻近专家;分析保证新增计算开销小,同时在不连续附近实现局部平滑。跨语言与视觉任务实验显示,该方法不仅使SMoE映射连续,还能提升实际性能。

原文摘要 · Abstract (English)

Sparse Mixture-of-Experts (SMoE) architectures are now widely deployed in state-of-the-art language and vision models, where conditional routing allows scaling to very large networks. However, this very Top-$k$ expert selection that enables conditional routing also renders the SMoE map inherently discontinuous. In the vicinity of these discontinuity surfaces, even inputs that are arbitrarily close may activate substantially different sets of experts resulting in significantly different outputs. In this work we give a rigorous geometric and stochastic analysis of these discontinuities. We first classify them by order, determined by the number of tied experts at a switching event. Using measure-theoretic slicing arguments, we establish asymptotic volume estimates for the thickened discontinuity surfaces, showing that lower-order discontinuity sets dominate, whereas higher-order ones occupy a vanishingly small relative volume. Next, modeling random perturbations in the input space via a diffusion process, we prove that the path eventually encounter a discontinuity, and moreover that the first hit almost surely occurs on an order-1 discontinuity with explicit finite-time probability bounds. We further derive occupation-time bounds that quantify the duration the random path spend in the neighborhoods of each discontinuity order. These theoretical results imply that inputs are more likely to lie near lower order discontinuities. Motivated by this insight, we propose a simple smoothing mechanism that can be directly applied to existing SMoEs, softly incorporating experts near discontinuities; our analysis guarantees that the added computational overhead remains small while providing localized smoothing near discontinuities, and experiments across language and vision tasks show that smoothing not only enforces continuity of the SMoE map but also enhances empirical performance.

专家混合模型稳定性随机分析平滑机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。