语言表述方式会影响视觉模型对图像的关注程度。
Tinted Frames: Question Framing Blinds Vision-Language Models
- 用语言框架作为探针,发现不同提问方式影响视觉注意力分配。
- 封闭式提问使模型关注图像信息减少30%以上,准确率下降显著。
- 提出轻量级提示调优方法,提升多框架下视觉理解一致性。
视觉语言模型(VLMs)常表现出对视觉输入的依赖不足,甚至在需要视觉推理的任务中也如此。本文揭示,这些模型存在选择性失明:其对图像的注意力会因语言表述方式而变化,即使不同表述所需视觉推理完全相同。通过分析视觉注意力分布,我们发现封闭式提问(如选择题、是/否)导致模型对图像上下文的关注度明显降低,任务相关区域注意力减弱,且更倾向于关注无信息量的词汇。这种注意力错配是造成准确率下降和跨框架不一致的主要原因。基于此机制洞察,我们提出一种轻量级提示调优方法,引入可学习标记以激发开放性提问中所见的稳健视觉锚定模式,在多种提问框架下均提升了视觉理解能力与性能表现。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have been shown to be blind, often underutilizing their visual inputs even on tasks that require visual reasoning. In this work, we demonstrate that VLMs are selectively blind. They modulate the amount of attention applied to visual inputs based on linguistic framing even when alternative framings demand identical visual reasoning. Using visual attention as a probe, we quantify how framing alters both the amount and distribution of attention over the image. Constrained framings, such as multiple choice and yes/no, induce substantially lower attention to image context compared to open-ended, reduce focus on task-relevant regions, and shift attention towards uninformative tokens. We further demonstrate that this attention misallocation is the principal cause of degraded accuracy and cross-framing inconsistency. Building on this mechanistic insight, we introduce a lightweight prompt-tuning method using learnable tokens that encourages the robust, visually grounded attention patterns observed in open-ended settings, improving visual grounding and improving performance across framings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。