研究单个物体如何携带场景信息,揭示视觉语言模型的上下文推理机制。
Contextual inference from single objects in Vision-Language models
- 通过遮蔽背景测试单个物体的上下文推断能力。
- 物体身份与场景类别预测准确率均高于随机水平。
- 不同模型对场景和大类上下文的依赖程度差异显著。
单个物体携带的场景上下文信息是人类场景感知中长期研究的问题,但视觉语言模型(VLMs)中的这一能力仍不清晰,直接影响模型鲁棒性。我们通过系统的行为与机制分析,探究从单个物体进行上下文推断的能力。在遮蔽背景条件下向VLMs呈现单个物体,测试其推断精细场景类别与粗粒度超类别(室内/室外)的能力。结果显示,单个物体在两个层级上均支持显著高于随机水平的推断,且性能受与人类场景分类一致的物体属性调节。物体身份、场景及超类别预测可部分解耦:某一层次的准确推断并不需要或保证另一层次的准确,且不同模型间耦合程度差异明显。机制层面,当背景被移除后仍保持稳定的物体表征,更有利于成功上下文推断。场景身份信息在整个网络中以图像标记形式编码,而超类别信息仅在后期出现或完全不存在。这些结果表明,VLMs中上下文推断的组织方式比准确率所反映的更为复杂,具有行为与机制上的双重特征。
原文摘要 · Abstract (English)
How much scene context a single object carries is a well-studied question in human scene perception, yet how this capacity is organized in vision-language models (VLMs) remains poorly understood, with direct implications for the robustness of these models. We investigate this question through a systematic behavioral and mechanistic analysis of contextual inference from single objects. Presenting VLMs with single objects on masked backgrounds, we probe their ability to infer both fine-grained scene category and coarse superordinate context (indoor vs. outdoor). We found that single objects support above-chance inference at both levels, with performance modulated by the same object properties that predict human scene categorization. Object identity, scene, and superordinate predictions are partially dissociable: accurate inference at one level neither requires nor guarantees accurate inference at the others, and the degree of coupling differs markedly across models. Mechanistically, object representations that remain stable when background context is removed are more predictive of successful contextual inference. Scene and superordinate schemas are grounded in fundamentally different ways: scene identity is encoded in image tokens throughout the network, while superordinate information emerges only late or not at all. Together, these results reveal that the organization of contextual inference in VLMs is more complex than accuracy alone suggests, with behavioral and mechanistic signatures
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。