arXiv:2609.06403cs.AI2026-09

提出路由有效秩,揭示MoE推理群体随时间分化与重聚的规律。

From Concentration to Differentiation and Back: Routing Effective Rank in MoE Reasoning Cohorts

论文配图:From Concentration to Differentiation and Back: Routing Effective Rank in MoE Reasoning Cohorts
图 1 · 摘自论文原文
  • 用专家路由相似性构建跨推理路径图,计算其熵有效维度
  • 发现98.5%的推理组呈现先集中、再分化、后重聚的三阶段轨迹
  • 可分解分析共模成分与残余谱,适合作为无标签诊断工具

测试时扩展生成推理路径群体,但缺乏对内部计算如何随推理演进重组的标准无标签描述。本文引入路由有效秩 deff,即基于MoE专家路由相似性构建的跨路径图的熵有效维度。在十种MoE配置和五个数学/科学基准上,deff呈现出可复现的低-高-低轨迹,3,105个模型-问题组中98.5%出现显著中间峰值:早期路由相似性集中,中期预算下分化最明显,后期重新集中,且峰值出现时间随架构与推理强度系统性变化。精确分解分离出群体共模质量与残余谱维度:共模重分配占轨迹约三分之二,残余谱贡献约四分之一,并在共模外仍保持显著变异。分解进一步定位行为特征:非一致组中,共模集中度上升强烈预示答案可恢复性;更高推理强度使峰值延迟2.59个八度(令牌预算翻倍),并一致延长高秩期,说明推理强度影响的是时机与持续时间而非峰值幅度。正确性对比表明,结构监控与答案选择可分离,路由有效秩作为可分解、无标签的诊断工具,提供了一种原理性的谱学视角,揭示MoE推理群体如何随推理时间分化与重聚。

原文摘要 · Abstract (English)

Test-time scaling produces cohorts of reasoning rollouts, yet there is no standard label-free account of how their internal computation reorganizes as inference unfolds. We introduce routing effective rank deff, the entropy-effective dimensionality of a cross-rollout graph built from MoE expert-routing similarity. Across ten MoE configurations and five math/science benchmarks, deff exhibits a reproducible low-high-low trajectory, with a prominent interior maximum in 98.5% of 3,105 model-question cohorts: routing similarity is concentrated early, maximally differentiated at intermediate budgets, and reconcentrated later, and the timing of this maximum varies systematically with architecture and reasoning effort. An exact decomposition separates cohort-wide common-mode mass from residual spectral dimensionality: common-mode reallocation accounts for about two thirds of the trajectory, while the residual spectrum contributes about one quarter and retains substantial variation beyond the common mode. The decomposition further localizes behavior: among non-unanimous cohorts, increases in common-mode concentration strongly predict same-answer recoverability, and higher reasoning effort delays the maximum by 2.59 octaves (doublings of the token budget) and consistently expands the high-rank period across all four tested architectures, locating the effort effect in timing and duration rather than peak amplitude. Correctness comparisons separate structural monitoring from answer selection, positioning routing effective rank as a decomposable, label-free diagnostic of cohort organization - a principled spectral lens on how MoE reasoning cohorts differentiate and reconcentrate over inference time.

MoE推理分析谱分析无标签诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。