arXiv:2510.01185cs.LG2025-10被引 1

用狄利克雷先验引导专家分工,提升混合专家模型性能

Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs

  • 通过狄利克雷先验损失直接调控路由分布,实现专家分工精细化
  • 在多模态模型上显著提升性能,优于传统微调与正则化方法
  • 无需人工干预即可注入领域先验,适用于各类概率输出模块

将预训练稠密模型高效转化为稀疏混合专家(MoE)可提升模型容量,但常因权重简单复制导致专家分工模糊。分析表明,即使采用常规正则化,转化后的MoE仍存在低置信度、弱区分性的路由行为,影响性能。本文提出狄利克雷先验塑造损失(DPSL),一种新型路由器正则化技术,通过匹配专家分配与目标狄利克雷先验,直接塑造路由概率分布。DPSL实现对专家平衡与专化的细粒度控制,支持编码如特定模态或任务聚焦等归纳偏置,无需人工干预;且该方法通用,适用于任何输出类别概率分布的模块。在基于Qwen2、Phi3、Llama3.2大语言模型骨干的视觉-语言上转换的MoE模型实验中,DPSL在标准基准测试中持续优于其他上采样策略与正则化方法,有效解决专家分工不佳问题,推动更自适应、高性能模型的发展。

原文摘要 · Abstract (English)

Upcycling pre-trained dense models into sparse Mixture-of-Experts (MoEs) efficiently increases model capacity but often suffers from poor expert specialization due to naive weight replication. Our analysis reveals that upcycled MoEs, even with conventional regularization, exhibit low-confidence, weakly differentiated routing, hindering performance. We introduce Dirichlet-Prior Shaping Loss (DPSL), a novel router regularization technique that directly shapes routing probability distributions by matching expert assignments to a target Dirichlet prior. DPSL offers fine-grained control over expert balance and specialization, and enables encoding of inductive biases such as encouraging experts to focus on specific modalities or tasks, without requiring manual intervention; notably, DPSL is a general tool applicable to any module that outputs categorical probability distributions, extending its utility beyond MoE training. Experiments on upcycled MoE vision-language models (with Qwen2, Phi3, Llama3.2 LLM backbones) show DPSL consistently outperforms upcycling strategies and regularization techniques across standard vision-language benchmarks, addressing the critical issue of poor specialization and fostering more adaptive, higher-performing models.

混合专家模型优化视觉语言路由机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。