arXiv:2605.25427cs.CVcs.AI2026-05

通过逐点指代让视觉语言模型解决多物体特征绑定问题

Binding Visual Features Point by Point

论文配图:Binding Visual Features Point by Point
图 1 · 摘自论文原文
  • 用文本显式指代物体位置,诱导模型产生内部视觉搜索机制
  • 微调后可消除绑定错误,实现组合泛化能力提升
  • 为视觉语言模型提供类人串行处理的可解释解决方案

尽管在标准基准上表现良好,视觉语言模型在处理多物体场景的任务中仍存在持续性失败,这些任务对人类而言相对简单。近期研究发现,这些失败可能源于在上下文中准确绑定物体特征的基本能力缺失,这一挑战在认知科学和神经科学中被称为“绑定问题”。人类视觉系统通过串行处理(一次关注一个物体)来解决此问题,避免其他物体干扰。最近工作提出“指代”——使用显式空间坐标指代物体——作为视觉语言模型的类似解决方案,并发现其能提升复杂多物体任务的表现。然而,尚不清楚该方法为何有效(即在机制或表征层面),以及与人类视觉串行处理的关系。本文研究此问题,发现通过文本学习指代会诱发内部视觉搜索流程,并刻画了支持该过程的机制。此外,指代行为可通过微调推广至新任务,消除绑定错误并实现组合泛化。结果证明,串行处理可为视觉语言模型解决绑定问题,如同在生物视觉中一样。

原文摘要 · Abstract (English)

Despite success on standard benchmarks, vision language models display persistent failures on tasks involving processing of multi-object scenes, including many tasks that are relatively easy for humans. Recent work has found that these failures may stem from a basic inability to accurately bind object features in-context, a challenge that is referred to as the "binding problem" in cognitive science and neuroscience. The human visual system is thought to solve this binding problem via serial processing, attending to individual objects one at a time so as to avoid interference from other objects. Recent work has proposed "pointing" -- the use of explicit spatial coordinates to refer to objects -- as an analogous solution for vision language models, and found that it improves performance on challenging multi-object tasks. However, it is unclear $\textit{why}$ (i.e., on a mechanistic or representational level) this approach improves performance, and how directly this relates to serial processing in human vision. Here, we investigate this question. We find that learning to point-via-text induces an internal visual search routine, and we characterize the mechanisms that support this procedure. We also find that pointing behavior can be generalized to new tasks via fine-tuning, and that doing so eliminates binding errors and enables compositional generalization. These results provide a proof-of-principle that serial processing can solve the binding problem for vision language models just as it does for biological vision.

视觉语言模型绑定问题串行处理指代机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。