arXiv:2606.28273cs.CL2026-06被引 1

揭示视觉语言模型在感知与知识冲突时的因果机制

Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models

论文配图:Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models
图 1 · 摘自论文原文
  • 通过激活修补和消融实验,定位关键注意力头
  • 仅2.5%-4.8%注意力头决定知识依赖性回答,占网络后半部分
  • 发现路由与写入头构成稀疏因果回路,适用于多种模型

视觉语言模型在视觉证据与记忆世界知识冲突时需做出判断。现有研究仅从行为层面描述该现象,缺乏组件级因果解释。本文结合三种粒度(残差流、注意力头、MLP子层)的激活修补、模型组件消融与机制分析,覆盖三类VLM。结果表明,视觉接地默认生效,而知识接地依赖于网络后半部分少数因果必要注意力头(占比2.5%-4.8%),这些头可使模型无视冲突视觉输入仍输出存储知识(如草莓为红色)。消融这些头后,在68%-96%的知识提示下预测由知识驱动转为视觉驱动,但仅改变0.8%-7.5%的视觉驱动预测,呈现显著不对称因果结构。这些头可分为路由头(调节信息流)与写入头(直接投影答案令牌至残差流),其结构在不同模型家族与规模中保持一致,揭示了视觉-知识冲突背后的稀疏因果回路。

原文摘要 · Abstract (English)

Vision-language models must reconcile visual evidence with memorized world knowledge when the two conflict. How they resolve this conflict shapes the reliability of multimodal systems, yet prior work characterizes it behaviorally without a component-level causal account. We combine activation patching across three granularities (residual stream, attention heads, and MLP sublayers) with model-component ablation studies and mechanistic analysis. Across three VLM families, we find that visual grounding emerges by default, whereas prior grounding depends on a small set of causally necessary attention heads (2.5-4.8%) concentrated in the second half of the network. These heads enable answers from stored world knowledge (e.g., "red" for a strawberry) despite conflicting visual input. Ablating them flips predictions from knowledge-grounded to visually grounded answers in 68-96% of cases under prior-knowledge prompts, but changes only 0.8-7.5% of visually grounded predictions, establishing an asymmetric causal structure. The identified heads decompose into routing heads, which modulate information flow, and writing heads, which directly project answer tokens into the residual stream. This structure is consistent across model families and scales, revealing a sparse causal circuit underlying perception-knowledge conflict in VLMs.

视觉语言模型因果机制知识冲突注意力头

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。