揭示视觉语言模型幻觉的双通路机制,可有效抑制错误生成。
Dual-Pathway Circuits of Object Hallucination in Vision-Language Models

- 提出双通路分析框架,区分正确推理与幻觉生成路径。
- 抑制幻觉通路可降低76%物体幻觉,且几乎不损失准确率。
- 该机制在不同模型间一致,可定向干预关系类幻觉。
视觉语言模型(VLMs)在跨模态推理中表现出色,但常产生输入图像中不存在的物体幻觉,影响其可靠性与可解释性。本文提出双通路电路分析框架,通过激活补丁技术在五种架构各异的VLMs中识别出支持正确预测的视觉定位通路与驱动错误输出的幻觉通路。引入条件通路分析(CPA)发现,定位组件在正确与幻觉样本中均保持强冗余性,但极性发生一致翻转:从支持真实答案转变为对幻觉答案的对齐。针对幻觉通路组件进行定向抑制,结果显示,扩大这些组件的规模可使物体幻觉减少高达76%,且精度损失极小;同时验证该电路仅能选择性迁移至关系类而非属性类幻觉。在POPE对抗性数据集和AMBER上的评估表明,所识别电路具有跨架构一致性、支持因果干预,并在不同幻觉类型间实现选择性迁移。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have demonstrated remarkable capabilities in bridging visual perception and natural language understanding, enabling a wide range of multimodal reasoning tasks. However, they often produce object hallucinations, describing content absent from the input image, which limits their reliability and interpretability. To address this limitation, we propose Dual-Pathway Circuit Analysis, a framework that identifies and characterizes hallucination-related circuits in VLMs for mechanistic understanding and causal probing. We first apply activation patching across five architecturally diverse VLMs to identify a visual grounding pathway that supports correct predictions and a hallucination pathway that drives erroneous outputs. We then introduce Conditional Pathway Analysis (CPA) to characterize pathway-level interactions, revealing that grounding components remain strongly redundant in both correct and hallucinating samples but undergo a consistent polarity flip, shifting from supporting the ground truth on correct samples to aligning with the hallucinated answer on erroneous ones. We further perform targeted suppression of hallucination-pathway components, showing that scaling these components reduces object hallucination by up to 76% with minimal accuracy cost, and validate that the same circuit selectively transfers to relational but not attribute hallucination. Evaluations on POPE-adversarial and AMBER show that the identified circuits are consistent across architectures, support causal intervention, and transfer selectively across hallucination types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。