揭示MoE专家功能解耦与表示重叠并存的几何规律
Geometric Asymmetry in MoE Specialization: Functional Decorrelation and Representational Overlap

- 通过雅可比-主成分-格拉斯曼框架分析专家功能与表示空间
- 专家功能高度解耦但表示子空间部分重叠,且路由稀疏性影响几何结构
- 适用于研究大模型条件计算机制,尤其适合关注专家分工的开发者
Mixture-of-Experts(MoE)架构通过稀疏路由实现可扩展容量,但专家专业化的几何结构仍不清晰。本文提出统一的雅可比-主成分-格拉斯曼分析框架,用于研究预训练MoE Transformer(Mistral、Qwen)中功能空间与表示空间的特性。结果发现:专家间功能解耦强烈(交叉雅可比对齐接近零),而其路由表示位于不同但部分重叠的子空间中,表明功能解耦与表示重叠可共存。控制路由实验显示,路由稀疏性是关键因素:top-k路由导致更强的功能分离和更大的子空间差异,而全软路由则产生更纠缠的专家结构。这些结果支持将MoE层视为在共享表示流形上于重叠子流形上实现局部解耦算子的几何解释,并为现代Transformer中的条件计算提供了通用诊断工具。
原文摘要 · Abstract (English)
Mixture-of-Experts (MoE) architectures achieve scalable capacity through sparse routing, yet the geometric structure of expert specialization remains poorly understood. We introduce a unified Jacobian-PCA-Grassmann framework for analyzing MoE layers in both function space and representation space. Across pretrained MoE Transformers (Mistral, Qwen), we find a consistent structural asymmetry: experts exhibit strong functional decorrelation (consistently low, near-zero cross-expert Jacobian alignment) while their routed representations occupy distinct but partially overlapping subspaces. This indicates that functional decorrelation and representation overlap coexist rather than coincide in MoE specialization. Controlled routing experiments further indicate that routing sparsity appears to be a key factor shaping this geometry: top-k routing induces sharper functional separation and larger subspace divergence, whereas fully soft routing yields more entangled expert structure. Together, these results suggest a geometric interpretation in which MoE layers may be viewed as implementing locally decorrelated operators over overlapping submanifolds on a shared representation manifold, and provide a general diagnostic framework for studying conditional computation in modern Transformer architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。