发现纹理驱动任务中模型依赖低频特征的捷径,剪枝后提升准确率与鲁棒性。
Low-Frequency Shortcuts in Texture-Driven Visual Learning

- 通过分析纹理驱动数据中的低频成分,揭示模型学习捷径机制。
- 剪除低频成分使域内准确率最高提升8%,对低频噪声鲁棒性增强40%。
- 适合关注模型泛化性、鲁棒性及纹理类视觉任务的研究者。
神经网络存在捷径学习问题,即在训练集上表现良好,但在分布内(ID)或分布外(OOD)测试集上泛化能力差。现有研究多基于形状驱动的标准基准,而众多实际应用为纹理驱动。本文首次分析纹理驱动场景下的捷径学习,发现其主要依赖少数低频成分(LFCs),尽管分类信息存在于高频细粒度细节中。剪除训练和测试集中的低频成分可消除捷径,实现更均衡的频谱行为,使域内准确率最高提升8%。低频捷径导致模型对分布外噪声极度敏感,最大准确率下降达70%。剪枝显著提升对低频扰动的鲁棒性(最高+40%),但带来高频扰动下的性能折损;频谱平衡改善整体泛化,但对高频特征的依赖降低高频频扰动下的表现。分布外准确率取决于这两者的交互作用。
原文摘要 · Abstract (English)
Neural networks suffer from shortcut learning, where learned features generalize well to the training set but not to in-distribution (ID) or out-of-distribution (OOD) test sets. Existing studies are all based on a few standard benchmarks, which are shape-driven. Numerous application domains, however, are texture-driven. In this work, we present shortcut learning analysis for texture-driven domains, and compare it with that of a standard benchmark. We show that texture-driven domains suffer from low-frequency shortcuts. They make the majority of their decisions based on a few low-frequency components (LFCs) with a skewed spectral behavior, despite that their classification information is in higher-frequency, fine-grained details. Pruning LFCs from training and test sets eliminates the shortcut and provides a more balanced spectral behavior, improving the ID accuracy by up to 8%. We show that low-frequency shortcuts make the models highly vulnerable to OOD corruptions, leading up to 70% accuracy drop compared to the ID accuracy. Pruning LFCs significantly improves robustness to low-frequency corruptions, by up to 40%, and introduces a trade-off for high-frequency corruptions; the balanced spectral behavior provides a better generalization performance, whereas the increased dependence on high-frequency features reduces it. OOD accuracy depends on the interaction between these two factors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。