arXiv:2602.08019cs.LGcs.AI2026-02综述被引 1

系统梳理稀疏专家混合模型的算法基础与应用前景

The Rise of Sparse Mixture-of-Experts: A Survey from Algorithmic Foundations to Decentralized Architectures and Vertical Domain Applications

  • 从路由机制与专家网络出发,剖析稀疏MoE的核心原理
  • 突破中心化架构局限,探索去中心化部署带来的效率与可扩展性提升
  • 覆盖横纵向应用场景,适合关注大模型高效训练的研究者

稀疏专家混合(Sparse Mixture of Experts, MoE)架构已成为在保持计算成本可控的前提下扩展深度学习模型参数量的有效方法。作为大语言模型的重要分支,MoE模型通过路由网络仅激活部分专家,实现稀疏条件计算,显著提升计算效率,为模型更大规模、更高性价比发展开辟了可行路径。它不仅在自然语言处理、计算机视觉及多模态等横向领域增强下游任务表现,也在垂直领域展现出广泛适用性。尽管MoE模型在多个领域广泛应用,但现有综述存在覆盖面不足或对关键方向探索不充分的问题。本文旨在填补这些空白:首先深入分析MoE的基础原理,包括路由网络与专家网络;其次拓展至去中心化范式,释放去中心化基础设施潜力,推动MoE开发民主化,实现更高可扩展性与成本效益;进一步探讨其在垂直领域的应用;最后指出当前挑战与未来研究方向。据我们所知,这是目前最全面的MoE领域综述,旨在为研究人员和实践者提供及时且有价值的参考。

原文摘要 · Abstract (English)

The sparse Mixture of Experts(MoE) architecture has evolved as a powerful approach for scaling deep learning models to more parameters with comparable computation cost. As an important branch of large language model(LLM), MoE model only activate a subset of experts based on a routing network. This sparse conditional computation mechanism significantly improves computational efficiency, paving a promising path for greater scalability and cost-efficiency. It not only enhance downstream applications such as natural language processing, computer vision, and multimodal in various horizontal domains, but also exhibit broad applicability across vertical domains. Despite the growing popularity and application of MoE models across various domains, there lacks a systematic exploration of recent advancements of MoE in many important fields. Existing surveys on MoE suffer from limitations such as lack coverage or none extensively exploration of key areas. This survey seeks to fill these gaps. In this paper, Firstly, we examine the foundational principles of MoE, with an in-depth exploration of its core components-the routing network and expert network. Subsequently, we extend beyond the centralized paradigm to the decentralized paradigm, which unlocks the immense untapped potential of decentralized infrastructure, enables democratization of MoE development for broader communities, and delivers greater scalability and cost-efficiency. Furthermore we focus on exploring its vertical domain applications. Finally, we also identify key challenges and promising future research directions. To the best of our knowledge, this survey is currently the most comprehensive review in the field of MoE. We aim for this article to serve as a valuable resource for both researchers and practitioners, enabling them to navigate and stay up-to-date with the latest advancements.

MoE大模型去中心化高效计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。