arXiv:2603.22278cs.CVcs.LG2026-03被引 2

揭示视觉语言模型中空间关系绑定的双机制,强调视觉编码器的核心作用。

The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models

  • 语言模型中间层处理对象间空间关系,但作用次要。
  • 视觉编码器全局编码物体布局,显著提升绑定性能。
  • 适用于改进多尺寸模型在复杂图像上的空间理解能力。

许多多模态任务(如图像描述和视觉问答)需要视觉语言模型(VLMs)将物体与其属性及空间关系绑定。然而,这种关联在VLMs中的计算位置与方式尚不明确。本文发现,VLMs依赖两种并行机制实现空间变量绑定:语言模型主干的中间层在对应物体的视觉标记上表示内容无关的空间关系,但此机制对模型预测影响较小;而主导的空间信息源自视觉编码器,其表征编码了物体布局,并被语言模型主干直接利用。值得注意的是,该空间信号在视觉标记间全局分布,延伸至物体区域之外的背景区域。我们证明,在所有图像标记上全局增强这些由视觉生成的空间表征,可显著提升不同规模模型在COCO数据集复杂自然图像上的空间变量绑定性能。结果揭示了VLMs内部空间绑定的计算机制,凸显视觉编码器的关键作用。

原文摘要 · Abstract (English)

Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Yet it remains unclear where and how such associations are computed within VLMs. In this work, we show that VLMs rely on two concurrent mechanisms to represent spatial variable binding. In the language model backbone, intermediate layers represent content-independent spatial relations on top of visual tokens corresponding to objects. However, this mechanism plays only a secondary role in shaping model predictions. Instead, the dominant source of spatial information originates in the vision encoder, whose representations encode the layout of objects and are directly exploited by the language model backbone. Notably, this spatial signal is distributed globally across visual tokens, extending beyond object regions into surrounding background areas. We show that enhancing these vision-derived spatial representations globally across all image tokens improves spatial variable binding performance across models of various sizes on complex natural images from the COCO datasets. Together, our results clarify how spatial variable binding is computed within VLMs and highlight the central role of vision encoders in enabling it.

视觉语言模型空间关系多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。