arXiv:2502.11453cs.LGcs.AI2025-02IJCAI综述被引 8

系统梳理多模态大模型中连接模块的设计演进与未来方向。

Connector-S: A Survey of Connectors in Multi-modal Large Language Models

  • 按原子操作与整体架构分类,构建连接模块的结构化体系。
  • 总结映射、压缩、专家混合等核心机制的技术进展。
  • 适合研究多模态融合与模型设计的学者参考。

随着多模态大语言模型(MLLMs)的快速发展,连接模块在弥合不同模态差异、提升模型性能方面发挥关键作用。然而,连接模块的设计与演进尚未得到系统分析,导致对其工作机制的理解不足,制约了更高效连接模块的发展。本文系统回顾了当前MLLM中连接模块的研究进展,提出一个结构化分类体系,将连接模块分为原子操作(如映射、压缩、专家混合)和整体设计(如多层、多编码器、多模态场景),并强调其技术贡献与演进路径。此外,讨论了若干有前景的研究前沿与挑战,包括高分辨率输入处理、动态压缩、引导信息选择、组合策略优化及可解释性问题。本综述旨在为研究人员提供基础参考与清晰路线图,助力下一代连接模块的设计与优化,以增强MLLM的性能与适应性。

原文摘要 · Abstract (English)

With the rapid advancements in multi-modal large language models (MLLMs), connectors play a pivotal role in bridging diverse modalities and enhancing model performance. However, the design and evolution of connectors have not been comprehensively analyzed, leaving gaps in understanding how these components function and hindering the development of more powerful connectors. In this survey, we systematically review the current progress of connectors in MLLMs and present a structured taxonomy that categorizes connectors into atomic operations (mapping, compression, mixture of experts) and holistic designs (multi-layer, multi-encoder, multi-modal scenarios), highlighting their technical contributions and advancements. Furthermore, we discuss several promising research frontiers and challenges, including high-resolution input, dynamic compression, guide information selection, combination strategy, and interpretability. This survey is intended to serve as a foundational reference and a clear roadmap for researchers, providing valuable insights into the design and optimization of next-generation connectors to enhance the performance and adaptability of MLLMs.

多模态连接模块大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。