arXiv:2603.06407cs.CV2026-03

揭示视觉Transformer如何决定图像前景背景,关键在后期层的凸性偏好。

Locating and Editing Figure-Ground Organization in Vision Transformers

  • 用合成飞镖图制造感知冲突,测试模型选择
  • 早期和中期层模糊,后期层突然偏好凸形完成
  • 发现特定注意力头(L0H9)是凸性偏见的源头

视觉Transformer需在局部几何线索与全局组织先验之间抉择前景-背景关系,导致典型感知模糊。本文通过基于合成飞镖形状的可控感知冲突,系统掩蔽可同时支持凹或凸补全的区域。实验表明,BEiT在竞争下稳定偏好凸形补全。通过逻辑归因将内部激活投影至离散视觉代码本空间,发现此偏好由变压器子结构中的可识别功能单元控制。具体而言,前景-背景组织在早期和中间层模糊,但在后期层突然明确。分解注意力头的直接效应后,发现头L0H9作为早期种子,引入微弱凸性偏见。缩小该头的权重,使感知冲突的分布质量沿连续决策边界移动,从而允许凹形证据主导补全。

原文摘要 · Abstract (English)

Vision Transformers must resolve figure-ground organization by choosing between completions driven by local geometric evidence and those favored by global organizational priors, giving rise to a characteristic perceptual ambiguity. We aim to locate where the canonical Gestalt prior convexity is realized within the internal components of BEiT. Using a controlled perceptual conflict based on synthetic shapes of darts, we systematically mask regions that equally admit either a concave completion or a convex completion. We show that BEiT reliably favors convex completion under this competition. Projecting internal activations into the model's discrete visual codebook space via logit attribution reveals that this preference is governed by identifiable functional units within transformer substructures. Specifically, we find that figure-ground organization is ambiguous through early and intermediate layers and resolves abruptly in later layers. By decomposing the direct effect of attention heads, we identify head L0H9 acting as an early seed, introducing a weak bias toward convexity. Downscaling this single attention head shifts the distributional mass of the perceptual conflict across a continuous decision boundary, allowing concave evidence to guide completion.

视觉Transformer前景背景注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。