让视觉语言模型的专家路由更懂模态,提升性能与效率
SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs

- 用动态软模态分数捕捉不同层的模态融合特征
- 在16个基准上平均提升0.9%多模态性能,通信开销降56.1%
- 适合部署在专家并行架构中的大模型开发者
混合专家(MoE)已成为大型视觉语言模型(VLMs)的主流骨干,但模态特定信号如何指导专家路由仍缺乏研究。现有路由策略或为手工设计,或无视模态,依赖理想化先验,忽略MoE-VLM中层间模态融合模式的变化,难以促进专家专业化。本文提出软模态引导的专家专业化(SMoES),包含捕捉层依赖融合模式的动态软模态分数、适配专家并行部署的专家分箱机制,以及促进模态专业化一致性的箱间互信息正则化。方法利用基于注意力或高斯统计的模态分数优化互信息正则化。在四个基于MoE的VLM和16个基准上的实验表明,该方法在有效性与效率上均有提升:多模态任务平均增益0.9%,语言任务增益4.2%;专家通信开销降低56.1%;在真实部署下吞吐量提升12.3%。结果验证了将路由与模态感知的专业化对齐可释放MoE-VLM的潜能。
原文摘要 · Abstract (English)
Mixture-of-Experts (MoE) has become a prevalent backbone for large vision-language models (VLMs), yet how modality-specific signals should guide expert routing remains under-explored. Existing routing strategies are either hand-crafted or modality-agnostic, relying on idealized priors that ignore the layer-dependent modality fusion patterns in MoE-VLMs and provide little guidance for expert specialization. We propose Soft Modality-guided Expert Specialization (SMoES), which consists of dynamic soft modality scores that capture layer-dependent fusion patterns, an expert binning mechanism aligned with expert-parallel deployment, and an inter-bin mutual information regularization that encourages coherent modality specialization. Our method leverages attention-based or Gaussian-statistics modality scores to optimize mutual information regularization. Experiments across four MoE-based VLMs and 16 benchmarks demonstrate improvement on both effectiveness and efficiency: 0.9% and 4.2% average gain on multimodal and language-only tasks, 56.1% reduction in EP communication overhead, and 12.3% throughput improvement under realistic deployment. These results validate that aligning routing with modality-aware expert specialization unlocks MoE-VLM capacity and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。