arXiv:2510.14349cs.CV2025-10被引 3

让多模态大模型更关注视觉信息,提升理解能力。

Vision-Centric Activation and Coordination for Multimodal Large Language Models

  • 引入视觉主导的激活与协调机制,融合多个视觉模型特征。
  • 在多个基准上显著提升多模态模型的视觉理解性能。
  • 适合需要强视觉推理能力的研究者和开发者使用。

多模态大语言模型(MLLMs)将视觉编码器提取的图像特征与大语言模型结合,展现出强大的理解能力。然而,主流MLLM仅通过文本下一个词预测进行监督,忽视了对分析能力至关重要的视觉中心信息。为此,我们提出VaCo,通过来自多个视觉基础模型(VFMs)的视觉中心激活与协调,优化MLLM的表征。VaCo引入视觉判别对齐,整合从VFMs中提取的任务感知感知特征,从而统一优化文本与视觉输出。具体地,我们在MLLM中引入可学习的模块化任务查询(MTQs)和视觉对齐层(VALs),在多个VFMs的监督下激活特定视觉信号;为协调不同VFMs间的表征冲突,设计了令牌网关掩码(TGM),限制多组MTQs之间的信息流动。大量实验表明,VaCo显著提升了多种MLLM在多个基准上的表现,展现出卓越的视觉理解能力。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) integrate image features from visual encoders with LLMs, demonstrating advanced comprehension capabilities. However, mainstream MLLMs are solely supervised by the next-token prediction of textual tokens, neglecting critical vision-centric information essential for analytical abilities. To track this dilemma, we introduce VaCo, which optimizes MLLM representations through Vision-Centric activation and Coordination from multiple vision foundation models (VFMs). VaCo introduces visual discriminative alignment to integrate task-aware perceptual features extracted from VFMs, thereby unifying the optimization of both textual and visual outputs in MLLMs. Specifically, we incorporate the learnable Modular Task Queries (MTQs) and Visual Alignment Layers (VALs) into MLLMs, activating specific visual signals under the supervision of diverse VFMs. To coordinate representation conflicts across VFMs, the crafted Token Gateway Mask (TGM) restricts the information flow among multiple groups of MTQs. Extensive experiments demonstrate that VaCo significantly improves the performance of different MLLMs on various benchmarks, showcasing its superior capabilities in visual comprehension.

多模态视觉理解模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。