arXiv:2512.19765cs.LGcs.AI2025-12AAAI被引 3

动态调整专家数量,让每个专家专精不同语义任务。

How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts

  • 根据语义漂移自动扩展专家池,避免盲目调参。
  • 在语言与视觉任务中均显著提升专家分工精度。
  • 适合需要高效、自适应专家分工的模型设计者。

寻找稀疏混合专家(SMoE)架构中最佳专家配置,以最大化专家间的语义区分度,是发挥MoE潜力的关键。然而,现有框架或过度依赖超参数调优,或忽视专家池规模调整时的语义角色多样性。本文提出面向自适应语义专精的混合专家框架MASS,引入两项关键改进:(i) 基于梯度的语义漂移检测器,在现有专家池无法覆盖数据全貌语义时触发精准专家扩展;(ii) 动态路由策略,依据令牌级路由置信度分布自适应调节专家使用。在受控合成实验中,MASS可靠收敛至成本-性能平衡点,并实现显著增强的语义专精。真实世界跨语言与视觉数据集上的实证结果表明,MASS持续优于多种强基线,展现出色的领域鲁棒性与专家分化能力。

原文摘要 · Abstract (English)

Finding the optimal configuration of Sparse Mixture-ofExperts (SMoE) that maximizes semantic differentiation among experts is essential for exploiting the full potential of MoE architectures. However, existing SMoE frameworks either heavily rely on hyperparameter tuning or overlook the importance of diversifying semantic roles across experts when adapting the expert pool size. We propose Mixture-of-Experts for Adaptive Semantic Specialization (MASS), a semanticaware MoE framework for adaptive expert expansion and dynamic routing. MASS introduces two key advancements: (i) a gradient-based semantic drift detector that prompts targeted expert expansion when the existing expert pool lacks capacity to capture the full semantic diversity of the data, and (ii) an integration of adaptive routing strategy that dynamically adjusts expert usage based on token-level routing confidence mass. We first demonstrate that MASS reliably converges to the point of optimal balance between cost-performance trade-off with notably improved sematic specialization in a highly controlled synthetic setup. Further empirical results on real-world datasets across language and vision domains show that MASS consistently outperforms a range of strong MoE baselines, demonstrating its domain robustness and enhanced expert specialization.

混合专家语义专精自适应扩展动态路由

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。