arXiv:2605.18498cs.LGcs.AI2026-05

提出首个系统性评估专家专精的框架,可精准优化大模型性能

DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs

  • 构建多领域基准与五项理论指标,分离功能专精与负载均衡
  • 发现不同模型有模块化或分布式专精模式,且专精程度影响性能提升
  • 通过诊断指标指导微调,仅用15%资源实现66%-94.48%性能提升

Mixture-of-Experts(MoE)模型中的专家专精机制仍不清晰,传统评估常将架构负载均衡与功能专精混淆。本文提出DBES,一个结合多领域基准与五项理论基础度量的综合诊断框架:路由专精、归一化有效秩、领域隔离、路由刚度评分和N-gram专精度量。关键发现表明,不同模型存在显著不同的专精范式:Qwen系列表现出高领域隔离的模块化专精,而DeepSeek与GLM则采用分布式协作模式。然而,我们强调专精是诊断维度,虽必要但非充分条件。最核心的是,干预实验证明这些度量可操作:在特定领域微调中利用DBES识别高专精路径,仅使用15%原始训练资源,即在专用领域实现66%至94.48%的性能提升。该工作首次提供独立于准确率的专家专精系统评估方法,为下一代MoE系统的设计与后训练优化提供关键洞见。

原文摘要 · Abstract (English)

Expert specialization in Mixture-of-Experts (MoE) models remains poorly understood, with traditional evaluations conflating architectural load-balancing with functional specialization. We introduce DBES, a comprehensive diagnostic framework combining a multi-domain benchmark with five theoretically grounded metrics: Routing Specialization, Normalized Effective Rank, Domain Isolation, Routing Stiffness Score, and N-gram Expertise measures. Critical findings demonstrate distinct specialization paradigms across models: Qwen-series exhibit modular specialization with high domain isolation, while DeepSeek and GLM employ distributed collaboration. However, we emphasize that specialization is a diagnostic dimension, necessary but not sufficient for downstream performance. Most crucially, interventional evidence validates the actionability of these metrics: by using DBES to identify high-specialization expert paths during domain-specific post-training, we achieved 66% to 94.48% improvement in specialized domains with only 15% of original training resources, demonstrating that these diagnostic tools can be converted into concrete optimization operators. This work provides the first systematic methodology for evaluating expert specialization independently of accuracy metrics, offering crucial insights for the design and post-training optimization of next-generation MoE systems.

MoE专家系统模型评估后训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。