视觉Transformer能从自然图像中学习到类似人类的图形-背景组织规律。
Vision Transformers Learn Gestalt-Like Figure-Ground Cues from Natural Images

- 用自然图像和人工刺激测试ViT对图形-背景线索的编码能力
- 模型在包围、凸性上表现稳定,零样本泛化至人工刺激
- 对称性仅在均匀色块中被编码,适合研究感知组织机制
人类视觉系统中的图形-背景组织依赖于包围、凸性和对称性等基于形状的线索。尽管这些线索在抽象刺激中已被广泛研究,但其在自然条件下的运作方式及其如何从自然场景统计中产生仍不明确。深度神经网络提供了新的研究路径:若模型使用与人类相同的图形-背景线索,则可为理解其底层机制提供可操作的实验手段。本研究评估了25个视觉变换器(ViTs),涵盖监督与自监督训练目标,通过线性探测器从中间图像块表示中预测图形-背景归属,使用自然图像及隔离单个线索的人工刺激。结果表明,ViTs稳健编码包围与凸性,且在自然图像上训练的探测器可零样本泛化至人工刺激;对称性则表现混合:均匀色块中有效,纹理区域中无效。整体说明,格式塔式图形-背景线索可从自然场景统计中学习,并使ViTs成为研究感知组织计算机制的理想模型系统。代码与数据见 https://github.com/mtangemann/mlvbench。
原文摘要 · Abstract (English)
Figure-ground organization in the human visual system relies on several shape-based cues, including surroundedness, convexity, and symmetry. While these cues have been extensively studied using abstract stimuli, little is known about how they operate under natural conditions or how they arise from the statistics of natural scenes. Deep neural networks offer a promising path forward: a model that relies on the same figure-ground cues as humans would provide tractable experimental access to the underlying mechanisms. In this study, we evaluate shape-based figure-ground organization in Vision Transformers (ViTs), for which prior work has demonstrated the emergence of object-based grouping. We test 25 ViTs spanning supervised and self-supervised training objectives, by fitting linear probes to predict figure-ground assignment from intermediate patch representations using both natural images and controlled artificial stimuli that isolate individual cues. Our results show that ViTs robustly encode surroundedness and convexity, and that probes trained on natural images generalize zero-shot to artificial stimuli across several models. For symmetry we observe mixed results: the cue is encoded for uniformly colored but not for textured regions. Taken together, our findings demonstrate that Gestalt-like figure-ground cues can be learned from natural scene statistics and position ViTs as a compelling model system for studying the computational mechanisms of perceptual organization. Code and data is available at https://github.com/mtangemann/mlvbench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。