系统梳理大模型主导的多模态融合方法,为未来模型设计提供框架指导。
Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques
- 以语言模型为中心,分类多模态输入融合架构与策略
- 分析125个模型,揭示融合机制与训练范式的演化趋势
- 适合研究多模态大模型融合技术的学者与工程师参考
多模态大语言模型(MLLMs)的快速发展重塑了人工智能格局。这些模型将预训练大语言模型与多种模态编码器结合,其融合需要对不同模态如何连接语言主干有系统理解。本综述提出一种以语言模型为中心的分析视角,考察将多样化模态输入转换并对齐至语言嵌入空间的方法,弥补现有文献的空白。我们基于三个关键维度构建分类框架:首先,分析模态融合的架构策略,包括具体融合机制与融合层级;其次,将表示学习技术分为联合或坐标式表示;第三,探讨训练范式,涵盖训练策略与目标函数。通过对2021至2025年间开发的125个MLLM进行分析,识别出该领域的新兴模式。本分类体系为研究者提供了当前集成技术的结构化概览,旨在指导基于预训练基础构建更稳健的多模态融合策略。
原文摘要 · Abstract (English)
The rapid progress of Multimodal Large Language Models(MLLMs) has transformed the AI landscape. These models combine pre-trained LLMs with various modality encoders. This integration requires a systematic understanding of how different modalities connect to the language backbone. Our survey presents an LLM-centric analysis of current approaches. We examine methods for transforming and aligning diverse modal inputs into the language embedding space. This addresses a significant gap in existing literature. We propose a classification framework for MLLMs based on three key dimensions. First, we examine architectural strategies for modality integration. This includes both the specific integration mechanisms and the fusion level. Second, we categorize representation learning techniques as either joint or coordinate representations. Third, we analyze training paradigms, including training strategies and objective functions. By examining 125 MLLMs developed between 2021 and 2025, we identify emerging patterns in the field. Our taxonomy provides researchers with a structured overview of current integration techniques. These insights aim to guide the development of more robust multimodal integration strategies for future models built on pre-trained foundations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。