arXiv:2609.05149cs.CLcs.CV2026-09

揭示视觉语言模型如何利用视觉信息做出决策,发现答案选项是关键的语义锚点。

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making

论文配图:From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making
图 1 · 摘自论文原文
  • 通过分层因果干预,追踪视频-文本注意力路径中的信息流动。
  • 视觉信息主要在处理选项时被整合,且名词比动词更关键。
  • 模型在时间推理上表现脆弱,可能受语言表达偏见影响。

视觉语言模型通常通过最终预测评估性能,但其决策是否基于视觉证据,需追溯视觉信息如何影响语言决策。为此,我们采用分层因果干预方法,在基于视频的生成式多选任务中研究跨模态信息流,聚焦空间、因果和时间推理。结果表明,视觉信息主要在模型处理候选答案选项时被整合,这些选项成为最终决策的主要文本锚点。进一步发现,名词在多模态丰富过程中起重要语义锚定作用,而动词在处理时间关系时更相关。此外,时间推理呈现特定模式:视觉语言模型难以重建跨视频帧的序列信息,但这种脆弱性也可能反映了与场景内事件关系描述相关的语言偏见。

原文摘要 · Abstract (English)

Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these decisions are grounded in visual evidence requires tracing how visual information contributes to language-based decisions. With this purpose in mind, we investigate cross-modal information flow in a video-based generative multiple-choice-like setting by applying a layer-wise causal intervention on video-text attention pathways. We target spatial, causal, and temporal visual reasoning. Our results show that visual information is mainly integrated while the model processes the candidate answer options, which serve as the primary textual grounding sites for the final decision. We further show that nouns play an important role as semantic anchors during multimodal enrichment, while verbs are more relevant when temporal relations are processed. Finally, we identify a distinct pattern in temporal reasoning, suggesting that VLMs struggle to reconstruct sequential information across video frames, but we remark that such fragility may also reflect linguistic biases associated with specific temporal expressions used for defining the relation between events within a scene.

视觉语言模型因果推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。