arXiv:2506.15871cs.CV2025-06被引 12

发现视觉语言模型中自发出现的空间索引符号机制,解决物体绑定难题。

Visual symbolic mechanisms: Emergent symbol processing in vision language models

  • 用空间独立的索引方式实现视觉对象绑定
  • 绑定错误可直接归因于该机制失效
  • 为减少模型绑定错误提供新思路

准确理解视觉场景需要将特征关联以表征独立物体。例如,区分一张含红方块与蓝圆的图像和一张含蓝方块与红圆的图像。近期研究发现语言模型通过一组内容无关的符号索引解决此‘绑定问题’,但视觉语言模型(VLMs)是否采用类似机制尚不明确。鉴于VLM在需绑定的任务上持续失败,本研究识别出一种此前未知的、在VLM中自发出现的符号机制,其通过内容无关的空间索引方案支持绑定。此外,我们发现当绑定出错时,可直接追溯至该机制的失效。这些结果揭示了支持VLM中类符号处理的内在机制,并提出降低绑定错误数量的可能路径。

原文摘要 · Abstract (English)

To accurately process a visual scene, observers must bind features together to represent individual objects. This capacity is necessary, for instance, to distinguish an image containing a red square and a blue circle from an image containing a blue square and a red circle. Recent work has found that language models solve this 'binding problem' via a set of symbol-like, content-independent indices, but it is unclear whether similar mechanisms are employed by Vision Language Models (VLMs). This question is especially relevant, given the persistent failures of VLMs on tasks that require binding. Here, we identify a previously unknown set of emergent symbolic mechanisms that support binding specifically in VLMs, via a content-independent, spatial indexing scheme. Moreover, we find that binding errors, when they occur, can be traced directly to failures in these mechanisms. Taken together, these results shed light on the mechanisms that support symbol-like processing in VLMs, and suggest possible avenues for reducing the number of binding failures exhibited by these models.

视觉语言模型符号机制绑定问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。