揭示视觉语言模型回答背后的真正视觉线索位置
Where To Look? : Causal Tracing of Vision Encoders in VLM

- 用因果追踪法分析模型决策依赖的视觉片段
- 关键视觉信息常在目标区域外,非局部定位
- 适合研究模型如何感知与使用视觉结构的研究者
视觉语言模型能精准描述图像,但其答案究竟依赖哪些视觉信息仍不明确。本文通过因果追踪发现,决定模型输出的关键视觉标记往往位于目标区域之外。在更大规模模型和不同扰动设置下,该现象普遍存在,表明强多模态性能并不等同于空间上局部化的因果表示。进一步实验表明,当外观线索被移除时,模型难以保持视觉结构理解,说明其依赖外观特征来把握视觉结构。这些结果揭示了‘感知’、‘使用’与‘推理’视觉结构之间的差距,并为研究现代视觉语言模型中视觉信息的转化、保留与运用提供了因果框架。
原文摘要 · Abstract (English)
Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actually drives their answers? In this work, we investigate this question through causal tracing, and we observe that highly causal vision tokens often lie outside the target region. Extending the analysis to larger vision-language models reveals a similar pattern across models and corruption settings, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations. We further investigate: can these models preserve visual structure when appearance cues are removed? and find that visual cues are exploited to understand visual structures. Together, our experiments expose a gap between seeing, using, and reasoning over visual structure, and provide a causal framework for studying how visual information is transformed, preserved, and ultimately used by modern vision-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。