arXiv:2604.07030cs.LG2026-04被引 1

小规模测试平台揭示MoE专家分工与路由行为规律

MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale

  • 构建真实数据混合的路由测试床,实现专家分工量化评估
  • 发现平衡覆盖范围是实现高利用率与专家专精的关键
  • 验证结论在35倍大模型上仍成立,适合模型优化研究者

稀疏混合专家(MoE)架构在前沿大语言模型中日益流行,但其路由复杂性带来了训练挑战。充分调动MoE模型参数需确保所有专家得到良好训练且具备非冗余的专业化能力。然而,由于缺乏成熟度量指标,且多数路由技术在小规模下表现相似,难以反映其大规模行为。为此,我们提出MoE路由测试床,通过将具有明显区分领域的数据组合与基准路由器结合,提供理想路由的明确上界,从而实现路由动态的清晰可视与可量化测量。实验表明,平衡路由覆盖范围是实现专家专精与高利用率的核心因素,并在35倍更大的模型中得到验证。

原文摘要 · Abstract (English)

Sparse Mixture-of-Experts (MoE) architectures are increasingly popular for frontier large language models (LLM) but they introduce training challenges due to routing complexity. Fully leveraging parameters of an MoE model requires all experts to be well-trained and to specialize in non-redundant ways. Assessing this, however, is complicated due to lack of established metrics and, importantly, many routing techniques exhibit similar performance at smaller sizes, which is often not reflective of their behavior at large scale. To address this challenge, we propose the MoE Routing Testbed, a setup that gives clearer visibility into routing dynamics at small scale while using realistic data. The testbed pairs a data mix with clearly distinguishable domains with a reference router that prescribes ideal routing based on these domains, providing a well-defined upper bound for comparison. This enables quantifiable measurement of expert specialization. To demonstrate the value of the testbed, we compare various MoE routing approaches and show that balancing scope is the crucial factor that allows specialization while maintaining high expert utilization. We confirm that this observation generalizes to models 35x larger.

MoE专家分工路由机制模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。