arXiv:2511.17793cs.CVcs.LG2025-11被引 3

通过注意力引导提升小模型视觉对齐,减少幻觉。

Attention Guided Alignment in Efficient Vision-Language Models

  • 用交错交叉注意力增强视觉定位能力
  • 在多个基准上显著降低物体幻觉率
  • 适合追求高效视觉语言理解的研究者

大型视觉语言模型依赖预训练视觉编码器与大语言模型之间的有效多模态对齐来整合视觉与文本信息。本文对高效视觉语言模型中的注意力模式进行了全面分析,发现基于拼接的架构常无法区分语义匹配与不匹配的图文对,这是导致模型产生物体幻觉的关键因素。为此,我们提出注意力引导的高效视觉语言模型(AGE-VLM),通过交错的交叉注意力层,将视觉能力注入预训练的小型语言模型中。该方法利用来自Segment Anything Model(SAM)的空间知识,强制模型关注正确的图像区域,显著减少幻觉现象。我们在多个以视觉为中心的基准上验证了该方法,结果优于或相当于现有高效VLM方法。研究为未来实现更优的视觉与语言理解提供了重要启示。

原文摘要 · Abstract (English)

Large Vision-Language Models (VLMs) rely on effective multimodal alignment between pre-trained vision encoders and Large Language Models (LLMs) to integrate visual and textual information. This paper presents a comprehensive analysis of attention patterns in efficient VLMs, revealing that concatenation-based architectures frequently fail to distinguish between semantically matching and non-matching image-text pairs. This is a key factor for object hallucination in these models. To address this, we introduce Attention-Guided Efficient Vision-Language Models (AGE-VLM), a novel framework that enhances visual grounding through interleaved cross-attention layers to instill vision capabilities in pretrained small language models. This enforces in VLM the ability "look" at the correct image regions by leveraging spatial knowledge distilled from the Segment Anything Model (SAM), significantly reducing hallucination. We validate our approach across different vision-centric benchmarks where our method is better or comparable to prior work on efficient VLMs. Our findings provide valuable insights for future research aimed at achieving enhanced visual and linguistic understanding in VLMs.

视觉对齐小模型幻觉抑制注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。