arXiv:2503.07137cs.LGcs.AI2025-03综述被引 126

系统梳理专家混合模型的算法、理论与应用,助力高效大模型发展

A Comprehensive Survey of Mixture-of-Experts: Algorithms, Theory, and Applications

  • 按门控机制、路由策略等拆解专家混合模型核心设计
  • 覆盖持续学习、多任务学习等前沿场景的算法创新
  • 适合关注大模型效率与可扩展性的研究者参考

人工智能在多个领域取得惊人进展,尤其得益于基础大模型的突破。这些大模型凭借海量训练数据,为众多下游任务提供通用解决方案。然而,随着数据日益多样复杂,大模型面临两大挑战:(1)计算资源消耗巨大,部署困难;(2)难以拟合异构复杂数据,限制实际可用性。专家混合(Mixture of Experts, MoE)模型近期受到广泛关注,通过动态选择并激活最相关的子模型处理输入数据,显著提升模型性能与效率,尤其在处理大规模多模态数据时表现优异。鉴于MoE在各领域的巨大潜力,亟需系统总结其最新进展。现有综述存在过时或关键领域缺失等问题,本文旨在填补空白。首先介绍MoE基本架构,包括门控函数、专家网络、路由机制、训练策略与系统设计;随后探讨其在持续学习、元学习、多任务学习及强化学习等重要机器学习范式中的算法设计;进一步总结理解MoE的理论研究,并回顾其在计算机视觉与自然语言处理中的应用;最后讨论有前景的未来研究方向。

原文摘要 · Abstract (English)

Artificial intelligence (AI) has achieved astonishing successes in many domains, especially with the recent breakthroughs in the development of foundational large models. These large models, leveraging their extensive training data, provide versatile solutions for a wide range of downstream tasks. However, as modern datasets become increasingly diverse and complex, the development of large AI models faces two major challenges: (1) the enormous consumption of computational resources and deployment difficulties, and (2) the difficulty in fitting heterogeneous and complex data, which limits the usability of the models. Mixture of Experts (MoE) models has recently attracted much attention in addressing these challenges, by dynamically selecting and activating the most relevant sub-models to process input data. It has been shown that MoEs can significantly improve model performance and efficiency with fewer resources, particularly excelling in handling large-scale, multimodal data. Given the tremendous potential MoE has demonstrated across various domains, it is urgent to provide a comprehensive summary of recent advancements of MoEs in many important fields. Existing surveys on MoE have their limitations, e.g., being outdated or lacking discussion on certain key areas, and we aim to address these gaps. In this paper, we first introduce the basic design of MoE, including gating functions, expert networks, routing mechanisms, training strategies, and system design. We then explore the algorithm design of MoE in important machine learning paradigms such as continual learning, meta-learning, multi-task learning, and reinforcement learning. Additionally, we summarize theoretical studies aimed at understanding MoE and review its applications in computer vision and natural language processing. Finally, we discuss promising future research directions.

专家混合大模型算法综述多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。