让专家路由更懂输入结构,提升模型稳定性和性能
STAR: Rethinking MoE Routing as Structure-Aware Subspace Learning

- 将路由视为结构感知子空间学习,动态追踪输入主成分
- 在语言与视觉任务中显著提升路由质量与下游表现
- 支持测试时更新子空间,增强分布偏移下的鲁棒性
混合专家(MoE)通过选择性地将输入路由到特定专家子集,高效扩展模型容量。然而,输入与专家之间的专业化依赖于路由机制是否真正感知输入结构。现实中,路由通常采用浅层线性投影,对输入表示的感知能力有限,常导致路由不稳定。本文提出STAR,将MoE路由重新构想为子空间学习问题,通过引入随时间演化的主子空间,利用广义赫布算法(GHA)追踪输入的主导结构,并与可学习路由结合。该设计使路由决策直接对齐输入结构,实现稳定的专家专业化。我们在受控合成设置及大规模语言与视觉任务上评估了STAR,结果表明其在路由质量与下游性能上持续优于强基准。此外,测试时可选的子空间更新进一步提升了路由在输入分布变化下的鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Mixture-of-Experts (MoE) scales model capacity efficiently by selectively routing inputs to a specialized subset of experts. However, input-expert specialization, the core motivation of MoE, critically depends on whether the router is actually aware of input structure. In practice, MoE routing is typically implemented as a shallow linear projection with limited awareness of input representation, which often leads to unstable routing. We propose STAR, a Structure Aware Routing that rethinks MoE routing as a subspace learning problem by augmenting standard learnable routing with an evolving principal subspace that tracks dominant input structure via Generalized Hebbian Algorithm (GHA). By aligning routing decisions directly with input structure, STAR enables stable expert specialization. We evaluate STAR on controlled synthetic setup and large-scale language and vision tasks, where it consistently improves routing quality and downstream performance over strong MoE baselines. Moreover, optional test-time subspace updates further enhance routing robustness and generalization under input distribution shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。