MoE模型通过软聚类方式降低函数局部曲率,提升表示多样性。
Mixture-of-Experts as Soft Clustering: A Dual Jacobian-PCA Spectral Geometry Perspective
- 用双重雅可比-主成分分析探测专家局部几何结构
- 软路由使专家间曲率更平缓、表示秩更高,拓扑重叠低
- 适合研究模型泛化性与专家协作机制的读者
Mixture-of-Experts(MoE)架构广泛用于提升效率和实现条件计算,但其对学习函数与表征几何的影响仍不清晰。本文从几何视角出发,将路由视为重叠专家局部图的软划分。提出双路径雅可比-主成分分析谱探针:通过雅可比奇异值谱分析局部函数几何,通过加权主成分分析路由后的隐藏状态揭示表征几何。在可控的MLP-MoE设置中,结合精确雅可比计算,对比密集、Top-k与全软路由在相同容量下的表现。结果表明:不同随机种子下,MoE路由均降低局部敏感性——专家局部雅可比的首项奇异值更小、谱衰减更快;加权主成分分析显示,专家局部表征在更多主方向上分布方差,有效秩更高。进一步发现专家雅可比间对齐度低,表明分解为低重叠的专家特定变换。路由锐度调节该效应:Top-k路由产生更集中、低秩结构,全软路由则生成更广、高秩表征。在三层Transformer与WikiText上的实验验证了自然语言任务中曲率下降,且Top-k路由跨专家对齐度更低。这些发现支持将MoE视为函数空间的软划分,能平滑局部曲率并重新分配表示方差,为专家扩展、幻觉抑制与集成多样性提供可检验预测。
原文摘要 · Abstract (English)
Mixture-of-Experts (MoE) architectures are widely used for efficiency and conditional computation, but their effect on the geometry of learned functions and representations remains poorly understood. We study MoEs through a geometric lens, interpreting routing as soft partitioning into overlapping expert-local charts. We introduce a Dual Jacobian-PCA spectral probe that analyzes local function geometry via Jacobian singular value spectra and representation geometry via weighted PCA of routed hidden states. Using a controlled MLP-MoE setting with exact Jacobian computation, we compare dense, Top-k, and fully soft routing under matched capacity. Across random seeds, MoE routing consistently reduces local sensitivity: expert-local Jacobians show smaller leading singular values and faster spectral decay than dense baselines. Weighted PCA reveals that expert-local representations distribute variance across more principal directions, indicating higher effective rank. We further observe low alignment among expert Jacobians, suggesting decomposition into low-overlap expert-specific transformations. Routing sharpness modulates these effects: Top-k routing yields more concentrated, lower-rank expert structure, while fully soft routing produces broader, higher-rank representations. Experiments on a 3-layer transformer with WikiText confirm curvature reduction on natural language and show lower cross-expert alignment for Top-k routing. These findings support interpreting MoEs as soft partitionings of function space that flatten local curvature while redistributing representation variance, yielding testable predictions for expert scaling, hallucination reduction, and ensemble diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。