arXiv:2608.22916cs.CL2026-08中稿 · EMNLP

揭示视觉语言模型中空间信息何时、如何传递到答案

Knowing Isn't Always Saying: When Do Spatial Encodings Reach Answers in Vision-Language Models?

论文配图:Knowing Isn't Always Saying: When Do Spatial Encodings Reach Answers in Vision-Language Models?
图 1 · 摘自论文原文
  • 用方向补丁技术追踪空间信息在模型各层的传输路径
  • 空间信息仅在中深层才影响最终答案,且受提示格式影响
  • 发现信息可延迟传递至最后输出阶段,解释了模型'知而未言'现象

视觉语言模型虽在隐藏状态中编码空间信息,却常不使用。我们通过方向补丁这一条件因果干预方法,在多层、多标记位置和多种提示格式下分析该信息何时抵达答案。基于已有编码证据构建的空间-ID方向,发现对答案逻辑值的因果影响仅在中深层出现。文本链式思考会抑制大多数模型中对象词的即时最大值传输,而视觉引导提示则保持通道开放。正向目标逻辑值增益可能低于最大值阈值,信息可在最终前缀标记或深层中的答案步骤重新出现。在研究的十种VLM中,这些局部效应形成可描述的传输模式。补充实验揭示这些模式随数据集、属性和编码强度的变化。结果将编码-接地差距重构为VLM中的条件传输问题。

原文摘要 · Abstract (English)

Vision-language models are known to encode spatial information in their hidden states, yet often fail to use it when answering. However, it remains unclear when and where this encoded information reaches the answer. We address this with direction patching, a class-conditioned causal intervention applied across layers, token positions, and prompt formats. Using spatial-ID directions constructed following prior encoding evidence, we find that causal influence on answer logits emerges only at mid-to-deep depths. Text chain-of-thought suppresses immediate object-word argmax-level transport in most models, while visually grounded prompts keep it open. Positive target-logit gain can remain below the argmax threshold, and transport can re-emerge at the final prefix token or at the answer step in deeper layers. Across the ten VLMs we study, these local effects form descriptive transport patterns. Complementary experiments characterize how these patterns shift across datasets, attributes, and encoding amplitudes. Together, these results reframe the encoding-grounding gap as a problem of conditional transport in VLMs.

视觉语言模型空间编码因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。