arXiv:2605.27431cs.LGcs.AI2026-05中稿 · IJCAI综述

用专家混合模型解决多模态学习中的效率与融合难题。

Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey

论文配图:Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey
图 1 · 摘自论文原文
  • 通过选择性激活专家,降低计算开销并减少模态冗余。
  • 融合多专家观点,增强跨模态表征对齐与交互能力。
  • 适配不完整或不平衡的多模态数据,灵活应对现实场景。

专家混合(Mixture-of-Experts, MoE)为多模态学习提供了天然兼容且可扩展的框架,展现出在多种模态和任务间的强适应性。尽管其应用日益广泛,但针对MoE如何应对多模态挑战的系统性综述仍较缺乏。现有综述多孤立评估多模态学习或MoE方法,忽视二者间独特互动。本文从三个核心视角回答关键问题: (1) MoE作为高效多模态引擎:通过解耦计算成本与参数增长,实现可扩展建模;(2) MoE作为多模态表征学习器:整合互补专家知识,丰富对齐与交互表示;(3) MoE作为多模态适配器:提供模块化机制,应对模态不平衡与缺失等不完美数据场景。通过广泛文献调研,我们识别出可解释路由、专家通信、模态融合及持续多模态学习等关键研究空白。本综述旨在为可解释、可持续的多模态MoE系统奠定基础。

原文摘要 · Abstract (English)

Mixture-of-Experts (MoE) presents a naturally compatible and scalable framework for multimodal learning, demonstrating strong adaptability across diverse modalities and tasks. Despite its growing success, a comprehensive and systematic review on the MoE metho addressing multimodal challenges remains lacking. Existing surveys tend to evaluate either multimodal learning or MoE independently from method taxonomy, overlooking the unique interplay between them. This survey fills that gap by answering a central question: \textit{How does MoE effectively resolve multimodal challenges?} We approach this from three key perspectives: (1) \textbf{MoE as an Efficient Multimodal Engine:} enabling scalable multimodal modeling by decoupling computational cost from parameter growth and mitigating modality redundancy through selective expert activation; (2) \textbf{MoE as a Multimodal Representation Learner:} integrating complementary multi-opinion expert knowledge to enrich alignment and interaction representations; and (3) \textbf{MoE as a Multimodal Adapter:} providing a modular and flexible mechanism to model imperfect data scenarios such as modality imbalance and missing modality. Through our extensive literature review, we identify critical research gaps, including interpretable routing, expert communication, modality integration, and lifelong multimodal learning. We position this survey as a foundation for future research toward interpretable and sustainable multimodal Mixture-of-Experts system.

多模态专家混合模型压缩表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。