arXiv:2510.08510cs.CVcs.AI2025-10被引 19

发现视觉模型中关键的高范数图像令牌,提升视觉推理能力。

To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models

  • 识别视觉编码器中重要性被忽视的高范数图像令牌
  • 这些令牌包含高层语义信息,显著提升模型理解力
  • 无需训练即可增强多类视觉推理任务表现

大型视觉语言模型(LVLM)融合视觉与文本信息,依赖视觉变换器(ViT)提取图像特征并生成图像标记序列,作为感知前端;而大语言模型(LLM)则处理这些标记进行高级推理。然而,哪些视觉标记对理解至关重要,其信号如何从ViT传递到LLM仍不清晰。现有研究多关注LLM中的注意力集中点(attention sinks),本文首次聚焦于ViT中的高范数视觉标记,称为ViT注意力集中点。分析表明,这些标记蕴含图像高层语义,对推理至关重要,但常被忽略。通过定性与定量分析其信息内容,提出无需训练和基于训练的方法以更有效利用这些标记。实验显示,显式使用这些标记可显著提升多种LVLM在多个视觉推理任务上的性能,揭示了ViT注意力集中点在增强视觉推理方面的巨大潜力。

原文摘要 · Abstract (English)

Large Vision Language Models (LVLMs) have recently emerged as powerful architectures capable of understanding and reasoning over both visual and textual information. These models typically rely on two key components: a Vision Transformer (ViT) and a Large Language Model (LLM). ViT encodes visual content into a sequence of image tokens and serves as the perceptual front-end -- the eyes of the model. In contrast, the LLM interprets these tokens to perform high-level reasoning, generates responses, and functions as the cognitive core -- the brain of the model. However, it remains unclear which visual tokens contribute most significantly to understanding and reasoning, and how effectively these signals are propagated from ViT to the LLM. While most existing works have focused on identifying attention sinks, low-semantic tokens receiving disproportionately high attention, within the LLM, we shift the focus to the vision encoder by identifying a class of high-norm visual tokens from ViT, referred to as ViT attention sinks -- a problem that has been rarely studied but is indeed very important for LVLMs. Our findings show that these ViT sinks encapsulate high-level semantic concepts from images, allowing the LLM to perform more effective understanding and reasoning. Despite their importance, these sink tokens are often overlooked in existing LVLM architectures. To explore their contribution, we present both qualitative and quantitative analyses of the information embedded in these sink tokens. We also propose both training-free and training-based approaches to better leverage how this information is interpreted by the LLM, and to what extent. By explicitly utilizing these tokens, we demonstrate substantial improvements across a range of LVLMs and visual reasoning tasks, highlighting the untapped potential of ViT attention sinks in enhancing visual reasoning.

视觉语言模型注意力机制视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。