arXiv:2607.16298cs.CV2026-07

用新指标发现模型更依赖纹理而非形状,且ViT比CNN更鲁棒。

Revisiting Shape and Texture Reliance with Category-Separability-Calibrated Suppression

论文配图:Revisiting Shape and Texture Reliance with Category-Separability-Calibrated Suppression
图 1 · 摘自论文原文
  • 提出语义退化指数SDI,以类别可分性为基准量化特征抑制影响。
  • 在相同退化程度下,五种CNN对纹理抑制的准确率下降比形状更严重。
  • ViT在图像分类和人脑视觉响应预测中均比CNN更抗干扰。

特征抑制评估通过削弱形状或纹理信息导致的准确率下降来推断模型依赖性。但这种下降同时受移除的类别相关性影响,且形状与纹理使用不同操作抑制,难以直接比较。本文提出语义退化指数(SDI),在由手工特征构建的固定判别空间中,量化抑制引起的类别可分性降低程度。在ImageNet16-like基准上,用高斯模糊(纹理抑制)与网格扭曲(形状抑制)在重叠退化范围内对比。当SDI值相同时,所有五种预训练卷积神经网络(CNNs)在高斯模糊下的准确率损失大于网格扭曲,表明其对纹理依赖更强,与以往在不匹配条件下的形状主导结论相反。所评估的视觉变换器(ViTs)在两种操作下普遍保持更高准确率。为验证该优势是否超越分类任务,还使用自然场景数据集中的干净与抑制图像评估固定脑编码模型。在两种抑制下,ViT特征的噪声上限归一化解释方差下降幅度均小于CNN。这些结果确立类别可分性作为解读抑制型特征依赖的关键参考,并表明CNN-ViT的鲁棒性差异延伸至预测人类视觉皮层反应的模型表征。

原文摘要 · Abstract (English)

Feature-suppression evaluations infer model reliance on shape or texture from the accuracy loss caused by attenuating each type of information. Such losses, however, conflate feature reliance with the amount of category-relevant information removed by the corresponding transformation. Because shape and texture are suppressed using different operators, their effects are not directly comparable. We introduce the Semantic Degradation Index (SDI), which quantifies the suppression-induced reduction in category separability relative to clean images in a fixed clean-reference discriminative space constructed from handcrafted features. On an ImageNet16-like benchmark, we use SDI to compare Gaussian blur for texture suppression with grid distortion for shape suppression over their overlapping degradation range. At comparable SDI values, all five evaluated ImageNet-trained convolutional neural networks (CNNs) retain less accuracy under Gaussian blur than under grid distortion. This results supports stronger texture than shape reliance under the evaluated operators, contrasting with the shape-dominant conclusion obtained from unmatched suppression conditions. The evaluated Vision Transformers (ViTs) also generally retain more accuracy than CNNs under both operators. To determine whether this advantage extends beyond classification, We evaluate fixed brain-encoding models using clean and suppressed images from the Natural Scenes Dataset. Under both operators, ViT features show smaller suppression-induced decreases in noise-ceiling-normalized explained variance than CNN features. These findings establish category separability as an important reference for interpreting suppression-based feature reliance and show that the CNN-ViT robustness difference extends to model representations predictive of human visual cortical responses.

特征依赖视觉模型类可分性ViT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。