arXiv:2608.00726cs.CV2026-08

用聚焦读取方法恢复视觉模型中的局部绑定信息

Foveated Probes Recover Localized Binding Information in Vision Foundation Models

  • 设计轻量聚焦读取,通过学习或问题引导的查询聚合图像块
  • 在干扰和反事实编辑下,聚焦读取恢复了接近理想读取的信号
  • 适合研究模型空间感知能力或改进视觉问答任务的学者

冻结的视觉基础模型通常通过单一全局图像嵌入进行评估,但这种接口可能混淆信息缺失与读取过程中的信息丢失。我们保持预训练视觉编码器冻结,仅改变其最终图像块令牌的读取方式。对比标准全局读取、轻量聚焦读取(使用学习或问题条件查询注意力池化图像块)以及可访问标注目标区域的理想读取。在三个局部绑定任务上评估:受干扰的合成颜色-形状绑定任务、无颜色的拥挤形状检测变体,以及基于GQA的自然图像任务,后者要求对同一图像中同类别对象的颜色进行成对询问。当合成目标单独出现时,全局读取表现接近完美,但在干扰和反事实目标编辑下严重退化;而聚焦读取则恢复了大部分理想读取可获取的信号。在GQA衍生任务中,不依赖问题的全局向量仅略优于仅依赖问题的先验,而问题条件聚焦显著提升配对局部颜色准确率。反事实干扰与信号比解释了合成任务的失败:全局池化稀释了局部标签变化证据,同时暴露于无关物体的干扰变化。结果表明,冻结视觉模型看似缺乏空间感知,实因全局嵌入接口所致,而非图像块令牌中无空间信息。

原文摘要 · Abstract (English)

Frozen vision foundation models are commonly evaluated through a single global image embedding, but this interface can conflate missing information with information lost at readout time. We study this distinction by keeping a pretrained vision encoder frozen and varying only the readout applied to its final patch tokens. We compare standard global readouts against a lightweight foveated readout, which attention-pools patch tokens using a learned or question-conditioned query, and against an oracle readout with access to the annotated target region. We evaluate these interfaces on three localized binding problems: a controlled synthetic color--shape binding task under clutter, a color-free crowded shape-detection variant, and a GQA-derived natural-image task where paired questions ask for the colors of different same-category objects in the same image. Global readouts perform near perfectly when the synthetic target appears alone, but collapse under clutter and counterfactual target edits, whereas the foveated readout recovers most of the oracle-accessible signal. On the GQA-derived task, question-independent global image vectors improve only modestly over question-only priors, while question-conditioned foveation substantially improves paired localized color accuracy. A counterfactual nuisance-to-signal ratio explains the synthetic failures: global pooling dilutes localized label-changing evidence while exposing the probe to nuisance variation from irrelevant objects. These results indicate that apparent spatial blindness in frozen vision models can arise from the global embedding interface rather than from an absence of spatial information in the frozen patch tokens.

视觉模型局部感知注意力机制模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。