发现视觉模型会伪造视觉理解,区分出两类幻觉机制。
Mirage Probes: How Vision Models Fake Visual Understanding

- 用对比探针框架识别模型内部的幻觉信号
- 两类幻觉:依赖语言先验和虚构视觉内容
- 需在表征层面干预才能真正实现视觉对齐
视觉-语言模型(VLMs)能在未提供图像的情况下自信且正确地回答图像相关问题,这种‘幻象’行为虚增了基准分数却无真实视觉依据。以往研究将其视为单一故障模式,我们提出应分为两种。通过名为Mirage Probes的对比探针框架,该框架将改写问题与匹配的幻象/非幻象标签配对于同一图像,我们发现幻象行为可在线性上从两个开源VLM的残差流、MLP、注意力后和注意力头等位置的内部激活中解码。我们证明朴素贝叶斯文本基线无法恢复此信号,排除了表面词汇混淆。跨基准分离模式结合新型先验利用指数(PHI)揭示两种不同范式:文本偏倚型,模型仅依赖语言先验而不调用视觉表征;虚假图像型,模型在隐空间构建错误视觉内容并似真作答。这一区分具有直接缓解意义:文本分布清洗可解决第一类,但无法触及第二类,因虚假图像幻象存在于模型的视觉表征中而非文本部分。真正可靠的视觉对齐需在表征层面实施干预。
原文摘要 · Abstract (English)
Vision-language models (VLMs) can answer image-based questions confidently, and often correctly, even when no image is provided. This mirage behavior inflates benchmark scores without reflecting visual grounding. Prior work treats this as a single failure mode. We argue it is two. Using Mirage Probes, a contrastive probing framework that pairs paraphrased question variants with matched mirage and non-mirage labels on the same image, we show that mirage behavior is linearly decodable from internal activations across residual stream, MLP, post-attention, and attention-head sites in two open-source VLMs. We demonstrate that a Naive Bayes text baseline cannot recover this signal, ruling out surface lexical confounds. Cross-benchmark separability patterns, together with a novel Prior Harnessing Index (PHI) measuring how much a model can answer from text alone, expose two distinct regimes: textual biases, where the model answers from language priors without engaging visual representations, and spurious images, where it constructs false visual content in latent space and answers as if grounded. The distinction has direct mitigation consequences: text-distribution cleaning can address the first regime but cannot reach the second, since spurious-image mirages live in the model's visual representations rather than its text. Faithful visual grounding will require interventions at the representational level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。