arXiv:2510.26560cs.LGstat.ML2025-10被引 1

发现深度网络中的捷径学习分布于全网,浅层存伪特征,深层丢真特征。

On Measuring Localization of Shortcuts in Deep Networks

  • 通过反事实训练量化各层对捷径导致准确率下降的贡献
  • 捷径学习非集中于某层,浅层存伪特征,深层丢核心特征
  • 揭示捷径分布差异,支持针对数据与架构定制化解法

捷径是训练中表现良好但泛化能力差的虚假规则,严重威胁深度网络可靠性。然而,捷径对特征表示的影响仍未被充分研究,阻碍了系统性缓解方法的设计。为此,本文研究深度模型中捷径的逐层定位问题。提出新颖实验设计,通过在干净与偏斜数据集上进行反事实训练,量化捷径诱导偏差对准确率下降的逐层贡献。在CIFAR-10、Waterbirds和CelebA数据集上,对VGG、ResNet、DeiT和ConvNeXt等架构进行分析,发现捷径学习并非集中在特定层,而是分布在整个网络中。不同网络部分扮演不同角色:浅层主要编码伪特征,深层则主要遗忘在干净数据上有预测性的核心特征。进一步分析了定位差异及其主要变化轴。最后,对逐层缓解策略的分析表明,设计通用方法难度大,支持采用数据与架构特异的方法。

原文摘要 · Abstract (English)

Shortcuts, spurious rules that perform well during training but fail to generalize, present a major challenge to the reliability of deep networks (Geirhos et al., 2020). However, the impact of shortcuts on feature representations remains understudied, obstructing the design of principled shortcut-mitigation methods. To overcome this limitation, we investigate the layer-wise localization of shortcuts in deep models. Our novel experiment design quantifies the layer-wise contribution to accuracy degradation caused by a shortcut-inducing skew by counterfactual training on clean and skewed datasets. We employ our design to study shortcuts on CIFAR-10, Waterbirds, and CelebA datasets across VGG, ResNet, DeiT, and ConvNeXt architectures. We find that shortcut learning is not localized in specific layers but distributed throughout the network. Different network parts play different roles in this process: shallow layers predominantly encode spurious features, while deeper layers predominantly forget core features that are predictive on clean data. We also analyze the differences in localization and describe its principal axes of variation. Finally, our analysis of layer-wise shortcut-mitigation strategies suggests the hardness of designing general methods, supporting dataset- and architecture-specific approaches instead.

捷径学习特征定位深度网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。