揭示视觉语言融合的深层机制,提升多模态模型推理能力
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
- 通过分层掩码分析发现视觉文本融合集中在特定层
- 提出无训练对比注意力框架,显著提升多模态推理性能
- 适合关注模型可解释性与融合优化的研究者
多模态大语言模型(MLLMs)在视觉-语言理解任务中取得显著进展,但其内部如何整合视觉与文本信息仍不清晰。本文对多种架构进行系统性的分层掩码分析,揭示了视觉-文本融合在MLLMs中的演化过程。结果显示,融合现象出现在若干特定层而非均匀分布,部分模型在输出生成前出现视觉信号的晚期‘重激活’现象。进一步分析分层注意力演化,发现无关区域存在持续的高注意力噪声,而文本对齐区域的注意力逐渐增强。基于这些洞察,我们提出一种无需训练的对比注意力框架,建模早期融合到最终层的注意力转换,突出有意义的关注转移。在多个MLLMs和基准上的实验验证了分析的有效性,并表明该方法能有效提升多模态推理性能。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language understanding, yet how they internally integrate visual and textual information remains poorly understood. To bridge this gap, we perform a systematic layer-wise masking analysis across multiple architectures, revealing how visual-text fusion evolves within MLLMs. The results show that fusion emerges at several specific layers rather than being uniformly distributed across the network, and certain models exhibit a late-stage "review" phenomenon where visual signals are reactivated before output generation. Besides, we further analyze layer-wise attention evolution and observe persistent high-attention noise on irrelevant regions, along with gradually increasing attention on text-aligned areas. Guided by these insights, we introduce a training-free contrastive attention framework that models the transformation between early fusion and final layers to highlight meaningful attention shifts. Extensive experiments across various MLLMs and benchmarks validate our analysis and demonstrate that the proposed approach improves multimodal reasoning performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。