揭示视觉语言模型依赖语言先验的内在机制与评估漏洞。
Diagnosing Visual Ignorance in Vision-Language Models

- 通过反事实层替换和逐层探针分析模型内部信息竞争
- 发现中层抑制视觉信息,后期强化文本偏见,导致视觉忽略
- 提出渐进模糊度量,验证多数题目在图像被破坏后仍可答对
视觉语言模型(VLMs)常依赖语言先验,生成看似自信但与视觉证据关联薄弱的答案。本文从机制与行为双重视角研究该问题。内部分析结合反事实层替换与监督式逐层MLP探针,追踪语言解码器中真实视觉语义与语言先验语义的竞争过程,发现多阶段瓶颈:中间层难以有效提取视觉信息,后期层进一步压制残存视觉信号以迎合文本空间偏见。外部评估引入基于多步高斯模糊的渐进视觉退化度量,识别出即使视觉内容被严重或完全破坏,答案仍保持不变的样本。在十二个视觉问答基准和三个代表性VLM上,大量示例在极端模糊下仍可作答,表明当前基准可能无意中奖励视觉无知。结果表明,语言先验依赖是系统性路由失败,影响模型内部结构与评估有效性。最后,提出未来研究路径,强调需设计结构隔离或反事实数据下的训练分布与评估协议,以强制跨模态真实对齐。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) frequently rely on language priors, producing confident answers that are weakly grounded in visual evidence. While this behavior is widely observed, its internal mechanisms and its impact on benchmark evaluation remain insufficiently understood. In this work, we study language-prior reliance from both mechanistic and behavioral perspectives. Internally, we combine counterfactual layer replacement with supervised layer-wise MLP probing to trace how ground-truth visual semantics and language-prior semantics compete across the language decoder. Our analysis reveals a multi-stage bottleneck: intermediate layers often fail to effectively retrieve visual information, while later layers can further suppress surviving visual signals in favor of text-space biases. Externally, we introduce a progressive visual decay metric based on multi-step Gaussian blurring, which identifies instances whose answers remain invariant even as visual content is increasingly destroyed. Across twelve visual question-answering benchmarks and three representative VLMs, we find that a substantial fraction of examples remain answerable under severe or total visual obfuscation, indicating that current benchmarks can inadvertently reward visual ignorance. These findings demonstrate that language-prior reliance is a systematic routing failure affecting both model internals and benchmark validity. Finally, we outline critical pathways for future research, highlighting the necessity of designing training distributions and evaluation protocols built on structurally isolated or counterfactual data to enforce genuine cross-modal grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。