arXiv:2507.00493cs.CVcs.AI2025-07NeurIPS被引 5

用形状谜题测试模型对整体结构的感知能力,发现顶尖模型能同时看懂纹理和布局。

Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models

  • 设计视觉谜题:打乱物体部件位置但保留纹理,检验模型是否认出整体形状
  • DINOv2等自监督模型在形状识别上表现最佳,远超传统模型
  • 高分模型依赖长程互动,且在中间深度出现从局部到全局编码的转变

人类可依据局部纹理与整体部件配置双重线索识别物体,而当前视觉模型多依赖纹理,导致特征脆弱、非组合性。现有研究将形状与纹理对立,忽略两者可并存,且无法衡量各自绝对质量。为此,本文提出配置形状得分(CSS),通过物体-字谜对(保留局部纹理但改变整体部件排列)评估模型对整体形状的感知能力。在86个卷积、变压器及混合模型中,CSS揭示了广泛分布的配置敏感度,其中完全自监督和语言对齐的变压器模型(如DINOv2、SigLIP2、EVA-CLIP)处于顶端。机制探针显示,高CSS模型依赖长程交互:半径控制注意力掩码使性能骤降,呈现独特的U型整合模式;表示相似性分析表明,在中等深度发生从局部到全局编码的转变。作为对照,BagNet模型仅随机表现,排除了‘边界黑客’策略的可能性。最终,我们证明配置形状得分还能预测其他依赖形状的任务表现。总体而言,真正鲁棒、泛化性强且类人化的视觉系统,可能不在于强制在形状与纹理之间做选择,而在于能无缝融合局部纹理与全局配置的架构与学习框架。

原文摘要 · Abstract (English)

Humans are able to recognize objects based on both local texture cues and the configuration of object parts, yet contemporary vision models primarily harvest local texture cues, yielding brittle, non-compositional features. Work on shape-vs-texture bias has pitted shape and texture representations in opposition, measuring shape relative to texture, ignoring the possibility that models (and humans) can simultaneously rely on both types of cues, and obscuring the absolute quality of both types of representation. We therefore recast shape evaluation as a matter of absolute configural competence, operationalized by the Configural Shape Score (CSS), which (i) measures the ability to recognize both images in Object-Anagram pairs that preserve local texture while permuting global part arrangement to depict different object categories. Across 86 convolutional, transformer, and hybrid models, CSS (ii) uncovers a broad spectrum of configural sensitivity with fully self-supervised and language-aligned transformers -- exemplified by DINOv2, SigLIP2 and EVA-CLIP -- occupying the top end of the CSS spectrum. Mechanistic probes reveal that (iii) high-CSS networks depend on long-range interactions: radius-controlled attention masks abolish performance showing a distinctive U-shaped integration profile, and representational-similarity analyses expose a mid-depth transition from local to global coding. A BagNet control remains at chance (iv), ruling out "border-hacking" strategies. Finally, (v) we show that configural shape score also predicts other shape-dependent evals. Overall, we propose that the path toward truly robust, generalizable, and human-like vision systems may not lie in forcing an artificial choice between shape and texture, but rather in architectural and learning frameworks that seamlessly integrate both local-texture and global configural shape.

视觉模型形状感知自监督学习架构设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。