arXiv:2509.06461cs.CVcs.AI2025-09被引 10

通过对比注意力提升视觉模型推理能力,无需训练

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning

  • 利用任务相关与通用查询的注意力对比,提取有效视觉信号
  • 在复杂场景下使模型推理准确率最高提升75%
  • 适合希望不训练即增强视觉模型性能的研究者

视觉语言模型(VLMs)在多种视觉任务中表现优异,但在复杂视觉环境中性能下降。现有增强方法多需额外训练、依赖外部分割工具或仅作用于粗粒度层级,忽视了VLM自身潜力。本文研究发现:(1) 视觉复杂度与注意力熵强相关,负向影响推理性能;(2) 注意力从浅层全局扫描逐步聚焦到深层集中,聚焦程度由视觉复杂度决定;(3) 理论证明,通用查询与任务特定查询的注意力对比可将视觉信号分解为语义成分与视觉噪声。基于此,提出无需训练的像素级注意力对比方法CARVE,有效提取任务相关视觉信号。大量实验表明,CARVE在开源模型上实现最高75%的性能提升,揭示了视觉复杂度与注意力机制的内在联系,为提升视觉推理提供了高效路径。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated remarkable success across diverse visual tasks, yet their performance degrades in complex visual environments. While existing enhancement approaches require additional training, rely on external segmentation tools, or operate at coarse-grained levels, they overlook the innate ability within VLMs. To bridge this gap, we investigate VLMs' attention patterns and discover that: (1) visual complexity strongly correlates with attention entropy, negatively impacting reasoning performance; (2) attention progressively refines from global scanning in shallow layers to focused convergence in deeper layers, with convergence degree determined by visual complexity. (3) Theoretically, we prove that the contrast of attention maps between general queries and task-specific queries enables the decomposition of visual signal into semantic signals and visual noise components. Building on these insights, we propose Contrastive Attention Refinement for Visual Enhancement (CARVE), a training-free method that extracts task-relevant visual signals through attention contrasting at the pixel level. Extensive experiments demonstrate that CARVE consistently enhances performance, achieving up to 75% improvement on open-source models. Our work provides critical insights into the interplay between visual complexity and attention mechanisms, offering an efficient pathway for improving visual reasoning with contrasting attention.

视觉推理注意力机制无训练增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。