arXiv:2506.00718cs.CVcs.AI2025-06被引 7

自监督模型能像人一样感知整体形状,关键在训练方式。

From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models

  • 用掩码自编码训练的视觉变压器可识别错觉轮廓和凸性偏好。
  • 在全局空间扰动测试中,自监督模型表现优于人类与监督模型。
  • 注意力机制非必需,但微调会破坏这种整体感知能力。

人类视觉通过格式塔原则(如闭合、邻近、图底分离)将局部线索整合为连贯的整体知觉,依赖于全局空间结构。本文研究现代视觉模型是否具备类似行为及其出现条件。发现采用掩码自编码(MAE)训练的视觉变压器(ViTs)表现出符合格式塔定律的激活模式,包括错觉轮廓补全、凸性偏好及动态图底分离。为探究其计算基础,提出新型测试基准DiSRT,评估模型对全局空间扰动的敏感度,同时保留局部纹理。结果表明,自监督模型(如MAE、CLIP)在该测试中表现优于监督基线,甚至超越人类表现。采用MAE训练的ConvNeXt也展现格式塔相容表征,说明此能力可在无注意力架构下产生。然而,分类微调会削弱该能力。受生物视觉启发,引入Top-K激活稀疏机制可恢复全局敏感性。研究揭示了促进或抑制格式塔感知的训练条件,并确立DiSRT作为跨模型全局结构敏感性诊断工具。

原文摘要 · Abstract (English)

Human vision organizes local cues into coherent global forms using Gestalt principles like closure, proximity, and figure-ground assignment -- functions reliant on global spatial structure. We investigate whether modern vision models show similar behaviors, and under what training conditions these emerge. We find that Vision Transformers (ViTs) trained with Masked Autoencoding (MAE) exhibit activation patterns consistent with Gestalt laws, including illusory contour completion, convexity preference, and dynamic figure-ground segregation. To probe the computational basis, we hypothesize that modeling global dependencies is necessary for Gestalt-like organization. We introduce the Distorted Spatial Relationship Testbench (DiSRT), which evaluates sensitivity to global spatial perturbations while preserving local textures. Using DiSRT, we show that self-supervised models (e.g., MAE, CLIP) outperform supervised baselines and sometimes even exceed human performance. ConvNeXt models trained with MAE also exhibit Gestalt-compatible representations, suggesting such sensitivity can arise without attention architectures. However, classification finetuning degrades this ability. Inspired by biological vision, we show that a Top-K activation sparsity mechanism can restore global sensitivity. Our findings identify training conditions that promote or suppress Gestalt-like perception and establish DiSRT as a diagnostic for global structure sensitivity across models.

视觉模型格式塔自监督感知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。