用图结构建模专家间互动,提升稀疏专家模型的鲁棒性。
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- 引入社交图结构模拟专家间交互,优化路由机制。
- 在4.2亿和7.4亿参数模型上验证有效性,提升抗数据污染能力。
- 轻量模块化设计,可无缝集成现有SMoE系统。
稀疏混合专家(SMoE)通过每样本仅激活少量参数,实现模型规模与计算成本解耦,具备显著扩展性。然而,传统SMoE在分布偏移下表现脆弱,易受数据污染影响。本文提出SymphonySMoE,引入社会图结构建模专家间交互,增强令牌路由机制,解决固有的鲁棒性问题。该方法轻量、模块化,可无缝集成至XMoE与通用语言模型等现有SMoE架构。理论分析与实证结果均表明其优于基线模型。在语言建模与视觉指令微调任务中广泛验证有效性,并成功拓展至4.2亿与7.4亿参数规模,展现出在大规模系统微调中的应用潜力。
原文摘要 · Abstract (English)
Sparse Mixture of Experts (SMoE) has emerged as a promising solution to achieving unparalleled scalability in deep learning by decoupling model parameter count from computational cost. By activating only a small subset of parameters per sample, SMoE enables significant growth in model capacity while maintaining efficiency. However, SMoE struggles to adapt to distributional shifts, leading to reduced robustness under data contamination. In this work, we introduce SymphonySMoE, a novel family of SMoE that introduces a social graph to model interactions among experts. This graph-based structure enhances the token routing process, addressing the robustness challenges that are inherent in conventional SMoE designs. SymphonySMoE is lightweight, modular, and integrates seamlessly with existing SMoE-based models such as the XMoE and the Generalist Language Model. We provide both theoretical analysis and empirical evidence demonstrating SymphonySMoE's advantages over baseline SMoE. Extensive experiments on language modeling and visual instruction tuning validate our method's effectiveness. We further highlight the scalability of SymphonySMoE to models with 4.2 and 7.4 billion parameters, showcasing its applicability in fine-tuning tasks for large-scale systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。