提出新路由机制,让专家分工更清晰且计算更均衡。
Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts
- 用概率混合模型划分输入空间,实现专家专注领域分离。
- 独立训练路由策略,避免任务目标干扰,提升稳定性。
- 在视觉语言任务中性能更强,专家使用更均衡。
稀疏混合专家(sMoE)已成为扩展大型视觉-语言模型的关键方法,在保持计算效率的同时提供强大容量。然而,现有路由机制通常基于相似性评分,难以有效捕捉输入结构,导致专家专业化与计算平衡之间存在权衡,限制了模型的可扩展性和性能。本文提出输入域感知的MoE框架,利用概率混合模型更好地划分输入空间。通过将路由概率建模为多个分布的混合,该方法使专家形成明确的专业化边界,同时实现均衡的使用。与传统方法不同,本路由机制独立于任务目标进行训练,确保优化稳定并实现明确的专家分配。在视觉-语言任务上的实验表明,该方法持续优于现有sMoE方法,取得更高的任务性能和更优的专家利用率平衡。
原文摘要 · Abstract (English)
Sparse Mixture of Experts (sMoE) has become a pivotal approach for scaling large vision-language models, offering substantial capacity while maintaining computational efficiency through dynamic, sparse activation of experts. However, existing routing mechanisms, typically based on similarity scoring, struggle to effectively capture the underlying input structure. This limitation leads to a trade-off between expert specialization and balanced computation, hindering both scalability and performance. We propose Input Domain Aware MoE, a novel routing framework that leverages a probabilistic mixture model to better partition the input space. By modeling routing probabilities as a mixture of distributions, our method enables experts to develop clear specialization boundaries while achieving balanced utilization. Unlike conventional approaches, our routing mechanism is trained independently of task-specific objectives, allowing for stable optimization and decisive expert assignments. Empirical results on vision-language tasks demonstrate that our method consistently outperforms existing sMoE approaches, achieving higher task performance and improved expert utilization balance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。