研究视觉变压器如何通过连续性原则绑定图像中的物体。
I Walk the Line: Examining the Role of Gestalt Continuity in Object Binding for Vision Transformers

- 用合成数据测试模型是否依赖连续性进行物体绑定
- 发现多个注意力头专门追踪物体的连续性线索
- 移除这些头会削弱模型的物体绑定能力,适合关注视觉机制的研究者
物体绑定是视觉认知中的基础过程,将低层感知特征整合为物体表征。尽管神经网络在这一任务上仍面临挑战,但近期研究表明预训练视觉模型中已出现绑定机制,能关联图像中属于同一物体的部分。本文探究视觉模型是否依赖格式塔连续性原则(而非相似性或邻近性)来实现物体绑定。通过合成数据集,我们证明绑定探测器对连续性具有敏感性,且在多种预训练视觉变压器中表现一致。进一步发现特定注意力头会追踪连续性,并在不同数据集间泛化。消融实验表明,这些头的移除会显著削弱模型生成物体绑定表征的能力。
原文摘要 · Abstract (English)
Object binding is a foundational process in visual cognition, during which low-level perceptual features are joined into object representations. Binding has been considered a fundamental challenge for neural networks, and a major milestone on the way to artificial models with flexible visual intelligence. Recently, several investigations have demonstrated evidence that binding mechanisms emerge in pretrained vision models, enabling them to associate portions of an image that contain an object. The question remains: how are these models binding objects together? In this work, we investigate whether vision models rely on the principle of Gestalt continuity to perform object binding, over and above other principles like similarity and proximity. Using synthetic datasets, we demonstrate that binding probes are sensitive to continuity across a wide range of pretrained vision transformers. Next, we uncover particular attention heads that track continuity, and show that these heads generalize across datasets. Finally, we ablate these attention heads, and show that they often contribute to producing representations that encode object binding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。