发现稀疏专家模型中路由状态具有共享几何结构,可统一建模提升预测精度。
Evidence for Shared Routing Geometry and Dynamics in Sparse Mixture-of-Experts

- 通过正交配准对齐各层路由控制空间,构建共享规范表示
- 共享动态模型在层间预测上保留79%~90%原始预测能力,$R^2=0.39$–$0.71$
- 该结构专用于路由决策,区别于普通隐藏态的平滑演化,适合模型压缩与优化
稀疏专家模型(MoE)在每一层使用独立参数化的路由器为每个令牌选择专家。已有研究发现,深层的路由决策可由浅层信号预测,表明路由并非完全独立。本文提供证据:不同层的路由相关状态共享共同的几何结构,但被层内坐标系掩盖。我们提取每层路由器的控制子空间,并通过广义正交普鲁斯特分析将其对齐至统一规范表示。对齐后,单一线性变换实现$R^2=0.39$–$0.71$,并保留79%–90%的原始逐层拟合预测能力,表明路由状态演化存在可复用的通用过程。进一步对比显示,残差表示跨层更易预测,而路由器控制状态更能忠实保持专家选择,区分了通用可预测性与路由特异性。最后测试表明,用预测的规范状态替代原生路由状态,仍能保持局部路由行为;学习状态演化相较简单持久化,在OLMoE上降低$ riangle ext{NLL}$达15.7%,在Phi上10层跨度降低6.2%。
原文摘要 · Abstract (English)
Sparse mixture-of-experts (MoE) models use an independently parameterized router at each sparse layer to select experts for every token. Prior work has shown that routing decisions across depth can often be predicted from earlier routing signals, suggesting that routing is not fully independent across layers. However, the structure behind this predictability remains unclear. In this work, we provide evidence that routing-relevant states across layers share a common geometric structure that is obscured by layer-specific coordinate systems. We isolate the control subspace of each router and align these spaces into a shared canonical representation using generalized orthogonal Procrustes analysis. After alignment, a single linear transition reaches $R^2=0.39$--$0.71$ and retains 79--90\% of the predictive power of separately fitted layer-specific dynamics, indicating that much of routing-state evolution follows a reusable process across depth. We then ask whether this shared dynamics is specific to routing or simply reflects the smooth evolution of hidden representations. A matched-rank comparison shows that residual representations are often easier to predict across layers, while router-control states preserve the model's expert choices much more faithfully. This separates generic cross-layer predictability from routing-specific information. Finally, we test whether the predicted canonical states remain meaningful when used in place of native routing states. The transported states preserve local routing behavior, while learned state evolution reduces $\Delta\mathrm{NLL}$ relative to simple persistence by 15.7\% on OLMoE and 6.2\% over a 10-router horizon on Phi.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。