揭示视觉语言模型中视觉信息的两种传递路径及其调控机制。
Pathways of Visual Information Flow in Vision-Language Models

- 通过因果补丁发现视觉信息经由直接与文本中介两条路径流动。
- 任务类型、数据分布和提示设计影响路径选择,且路径可灵活切换。
- 模型在路径失效时能启用备用路径,揭示其内部机制的鲁棒性。
我们研究视觉语言模型(VLMs)中视觉信息的传递路径。基于可控合成与自然数据集的因果补丁实验,发现模型解决视觉任务依赖两种不同路径:一是直接路径,视觉信息保留在图像标记中,并由后层最终标记读取;二是文本中介路径,视觉信息先转移至查询标记,再由最终标记读取。在三个视觉任务中,路径选择具有任务依赖性,且数据分布与提示设计亦可调节路径使用。进一步通过注意力敲除与损坏输入补丁分析发现,当常规路径被破坏时,模型可依赖文本中介路径作为后备。该行为统一了先前研究发现,表明消融干预可揭示模型潜在能力而非常态行为。结果为VLMs中视觉信息流提供了机制性解释,并凸显其内部机制在干预下的灵活性。
原文摘要 · Abstract (English)
We study how visual information is routed in vision-language models (VLMs). Using causal patching on controlled synthetic and natural datasets, we find that models rely on two distinct pathways to solve visual tasks: A direct pathway, where visual information is retained in image token representations and read out by the final token at later layers, and a text-mediated pathway, where visual information is first transferred to the query tokens and then read out by the final token. Across three visual tasks, we show that pathway selection is task-dependent, and that data distribution and prompt design can also modulate which pathway is used to solve the image-based query. Moreover, using attention knockouts and corrupted-input patching, we find that these pathways are flexible, under certain interventions, models can rely on the text-mediated pathway as a fallback when the usual pathway is ablated. This behavior unifies findings in prior work and shows that ablation-based interventions can reveal what models could do rather than what they normally do. Together, our results provide a mechanistic characterization of visual information flow in VLMs and highlight the flexibility of their internal mechanisms under intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。