arXiv:2602.17798cs.LG2026-02被引 1

提出基于子空间流形的MoE路由机制,实现可调稀疏性与负载均衡。

Grassmannian Mixture-of-Experts: Concentration-Controlled Routing on Subspace Manifolds

  • 在子空间流形上用矩阵宾汉分布控制路由熵,单一参数调节稀疏性
  • 0%专家坍塌,负载均衡提升15–30%,支持训练后调稀疏性
  • 专家浓度差异反映语言专长,路由行为可解释,适合需要稳定路由的研究

Mixture-of-Experts 模型依赖学习的路由器分配令牌到专家,但标准 softmax 路由无法原则性地控制稀疏性与利用率之间的权衡。我们提出 Grassmannian MoE(GrMoE),一种在子空间流形上的路由框架,其门控权重来源于矩阵宾汉分布的集中参数。该构造产生一个单一、可解释的控制参数——集中矩阵 Λ——连续调控路由熵,将离散的 top-k 选择替换为平滑且几何合理的稀疏机制。我们进一步开发了一种近似变分推断方法,用于后验路由分布,实现不确定性感知的专家分配,自然抵抗专家坍塌。我们正式证明了集中谱与路由熵、期望 top-k 质量及专家坍塌指数边界之间的紧密关系,建立了首个集中控制稀疏性的形式理论。在合成路由任务中,具有8个专家的3.5亿参数模型、16个专家的13亿参数模型和32个专家的27亿参数模型,均实现0%路由坍塌,困惑度相当或更优,负载均衡提升15–30%,且浓度与有效稀疏性呈平滑单调关系,支持无需重训练的后处理稀疏性调节。逐令牌分析显示,专家学习到异质的浓度值,与语言专业化相关,提供可解释的路由行为。

原文摘要 · Abstract (English)

Mixture-of-Experts models rely on learned routers to assign tokens to experts, yet standard softmax gating provides no principled mechanism to control the tradeoff between sparsity and utilization. We propose Grassmannian MoE (GrMoE), a routing framework that operates on the Grassmannian manifold of subspaces, where gating weights arise from the concentration parameters of Matrix Bingham distributions. This construction yields a single, interpretable knob -- the concentration matrix $Λ$ -- that continuously controls routing entropy, replacing discrete top-$k$ selection with a smooth, geometrically principled sparsity mechanism. We further develop an amortized variational inference procedure for posterior routing distributions, enabling uncertainty-aware expert assignment that naturally resists expert collapse. We formally prove tight bounds relating the Bingham concentration spectrum to routing entropy, expected top-$k$ mass, and an exponential bound on expert collapse, establishing the first formal theory of concentration-controlled sparsity. On synthetic routing tasks, a 350M-parameter MoE language model with 8 experts, a 1.3B-parameter model with 16 experts, and a 2.7B-parameter model with 32 experts, GrMoE achieves 0\% routing collapse across all seeds, comparable or better perplexity with 15--30\% improved load balance, and a smooth monotonic relationship between concentration and effective sparsity that enables post-hoc sparsity tuning without retraining. Token-level analysis reveals that experts learn heterogeneous concentration values that correlate with linguistic specialization, providing interpretable routing behavior.

MoE稀疏路由流形学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。