arXiv:2410.06765cs.CLcs.CV2024-10EMNLP被引 13

对比了视觉连接器的保留与压缩效果,发现细粒度任务用保留型更准,粗粒度任务用压缩型更快。

To Preserve or To Compress: An In-Depth Study of Connector Selection in Multimodal Large Language Models

  • 将连接器分为保留细节和压缩特征两类,统一评估不同任务表现。
  • 保留型连接器在细粒度感知任务上显著更优,压缩型在粗粒度和推理任务中速度更快且性能相当。
  • 研究结果可指导多模态大模型架构设计,适合关注效率与精度权衡的开发者。

近年来,多模态大语言模型(MLLMs)受到产业界与学术界的广泛关注。然而,在构建MLLM架构方面仍存在争议,尤其在针对不同粒度感知任务时如何选择合适的连接器。本文系统研究了连接器对MLLM性能的影响。具体地,我们将连接器分为特征保留型与特征压缩型,并基于统一标准,将MMBench、MME和SEED-Bench三个综合性基准中的子任务划分为粗粒度感知、细粒度感知和推理三类,进行评估。研究发现,特征保留型连接器在细粒度感知任务中表现优异,因其能有效保留详细视觉信息;而特征压缩型连接器虽在细粒度任务中表现较差,但在粗粒度感知和推理任务中具有显著加速优势且性能相当。这些发现对指导MLLM架构设计及优化具有重要意义。

原文摘要 · Abstract (English)

In recent years, multimodal large language models (MLLMs) have garnered significant attention from both industry and academia. However, there is still considerable debate on constructing MLLM architectures, particularly regarding the selection of appropriate connectors for perception tasks of varying granularities. This paper systematically investigates the impact of connectors on MLLM performance. Specifically, we classify connectors into feature-preserving and feature-compressing types. Utilizing a unified classification standard, we categorize sub-tasks from three comprehensive benchmarks, MMBench, MME, and SEED-Bench, into three task types: coarse-grained perception, fine-grained perception, and reasoning, and evaluate the performance. Our findings reveal that feature-preserving connectors excel in \emph{fine-grained perception} tasks due to their ability to retain detailed visual information. In contrast, feature-compressing connectors, while less effective in fine-grained perception tasks, offer significant speed advantages and perform comparably in \emph{coarse-grained perception} and \emph{reasoning} tasks. These insights are crucial for guiding MLLM architecture design and advancing the optimization of MLLM architectures.

多模态模型压缩连接器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。