让微调模块像专家路由一样动态选择,提升大模型效率
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
- 将路由机制引入微调模块,与MoE架构匹配
- 在OLMoE和Mixtral模型上实现更高性能与效率
- 为不同场景提供可复用的微调配置建议
混合专家(MoE)通过专家间的动态路由机制获益,但现有参数高效微调(PEFT)策略未能充分利用这一特性。为此,我们探究适应模块是否应引入路由机制以契合MoE的多专家架构。分析了在MoE语言模型上应用PEFT时核心组件的动态行为,并研究不同路由策略对适应效果的影响。在OLMoE-1B-7B和Mixtral-8x7B模型上,针对多种常识推理与数学推理任务的大量实验验证了所提出路由方法在性能与效率上的优势。我们识别出不同场景下的最优配置,并提供了实证分析与实用洞察,助力更优的PEFT与MoE应用。
原文摘要 · Abstract (English)
Mixture-of-Experts (MoE) benefits from a dynamic routing mechanism among their specialized experts, which existing Parameter- Efficient Fine-Tuning (PEFT) strategies fail to leverage. This motivates us to investigate whether adaptation modules themselves should incorporate routing mechanisms to align with MoE's multi-expert architecture. We analyze dynamics of core components when applying PEFT to MoE language models and examine how different routing strategies affect adaptation effectiveness. Extensive experiments adapting OLMoE-1B-7B and Mixtral-8x7B on various commonsense and math reasoning tasks validate the performance and efficiency of our routed approach. We identify the optimal configurations for different scenarios and provide empirical analyses with practical insights to facilitate better PEFT and MoE applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。