揭示视觉语言模型在视觉与知识冲突时的决策机制
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
- 构建反事实数据集,探测视觉与常识知识的冲突
- 发现少数注意力头主导冲突解决,干预可控制模型倾向
- 定位视觉优先区域更精准,优于传统梯度方法
视觉语言模型(VLM)融合视觉与文本信息执行复杂任务,但其内部知识与外部视觉输入冲突时易产生幻觉和不可靠预测。本文通过引入WHOOPS-AHA!数据集——一组故意违背常识的多模态反事实问题,研究VLM解决跨模态冲突的机制。通过日志值分析,识别出少数关键注意力头负责协调冲突。干预这些头可引导模型偏向内部参数化知识或视觉信息。结果显示,这些注意力头的激活模式能有效定位影响视觉主导决策的图像区域,相较于基于梯度的方法具有更精确的归因能力。
原文摘要 · Abstract (English)
Vision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks. However, conflicts between their internal knowledge and external visual input can lead to hallucinations and unreliable predictions. In this work, we investigate the mechanisms that VLMs use to resolve cross-modal conflicts by introducing WHOOPS-AHA!, a dataset of multimodal counterfactual queries that deliberately contradict internal commonsense knowledge. Through logit inspection, we identify a small set of attention heads that mediate this conflict. By intervening in these heads, we can steer the model towards its internal parametric knowledge or the visual information. Our results show that attention patterns on these heads effectively locate image regions that influence visual overrides, providing a more precise attribution compared to gradient-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。