arXiv:2510.00475cs.LGcs.CV2025-10

提出新指标诊断持续学习中因捷径引发的僵化问题

Diagnosing Shortcut-Induced Rigidity in Continual Learning: The Einstellung Rigidity Index (ERI)

  • 用三个可解释维度分离真实迁移与干扰线索导致的性能
  • 发现多数方法虽加速适应但最终准确率下降,说明捷径有害
  • 适合研究模型泛化性与持续学习鲁棒性的学者参考

深度神经网络常利用无因果意义的偶然相关特征(即捷径特征),这些特征在分布偏移下会降低模型可靠性。在持续学习(CL)中,这种捷径依赖可能持续并加剧:先前任务继承的权重会偏向重复使用易于满足旧标签的特征,类似认知中的‘定势效应’。与灾难性遗忘不同,捷径引起的僵化会阻碍新知识的学习。本文提出Einstellung僵化指数(ERI),通过三个可解释维度诊断:(i) 适应延迟(AD)、(ii) 性能缺陷(PD)、(iii) 相对次优特征依赖度(SFR_rel)。在两阶段CIFAR-100基准上,第2阶段故意引入品红色补丁作为虚假线索。评估了朴素微调(SGD)、在线弹性权重巩固(EWC_on)、Dark经验回放(DER++)、梯度投影记忆(GPM)和深度生成回放(DGR)。结果显示,大多数方法比从头训练(Scratch-T2)更早达到准确率阈值(负AD),但在含补丁类别上最终准确率略低(正PD)。屏蔽补丁后,CL方法准确率提升,而Scratch-T2略有下降,体现负SFR_rel。表明在此设置中,补丁是干扰项而非有效捷径。

原文摘要 · Abstract (English)

Deep neural networks frequently exploit shortcut features, defined as incidental correlations between inputs and labels without causal meaning. Shortcut features undermine robustness and reduce reliability under distribution shifts. In continual learning (CL), the consequences of shortcut exploitation can persist and intensify: weights inherited from earlier tasks bias representation reuse toward whatever features most easily satisfied prior labels, mirroring the cognitive Einstellung effect, a phenomenon where past habits block optimal solutions. Whereas catastrophic forgetting erodes past skills, shortcut-induced rigidity throttles the acquisition of new ones. We introduce the Einstellung Rigidity Index (ERI), a compact diagnostic that disentangles genuine transfer from cue-inflated performance using three interpretable facets: (i) Adaptation Delay (AD), (ii) Performance Deficit (PD), and (iii) Relative Suboptimal Feature Reliance (SFR_rel). On a two-phase CIFAR-100 CL benchmark with a deliberately spurious magenta patch in Phase 2, we evaluate Naive fine-tuning (SGD), online Elastic Weight Consolidation (EWC_on), Dark Experience Replay (DER++), Gradient Projection Memory (GPM), and Deep Generative Replay (DGR). Across these continual learning methods, we observe that CL methods reach accuracy thresholds earlier than a Scratch-T2 baseline (negative AD) but achieve slightly lower final accuracy on patched shortcut classes (positive PD). Masking the patch improves accuracy for CL methods while slightly reducing Scratch-T2, yielding negative SFR_rel. This pattern indicates the patch acted as a distractor for CL models in this setting rather than a helpful shortcut.

持续学习捷径学习模型诊断泛化性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。