arXiv:2601.22012cs.LG2026-01被引 3

从特征编码角度解释持续学习中的灾难性遗忘,揭示深度模型的脆弱性。

Putting a Face to Forgetting: Continual Learning meets Mechanistic Interpretability

  • 将遗忘建模为特征编码的几何变换,揭示容量压缩与读出破坏机制
  • 实验验证深度加剧遗忘,且在玩具模型中可识别最优/最差场景
  • 适用于视觉变压器等实际模型,帮助理解特征级动态变化

持续学习中的灾难性遗忘通常仅在性能或最后一层表示层面衡量,忽略了底层机制。本文提出一种机制性框架,将灾难性遗忘几何化为单个特征编码的变换结果。这些变换可能导致遗忘,通过减少特征分配容量或干扰下游计算对特征的读取。通过对一个可解析的简化模型进行形式化分析,我们明确了最佳与最差情况。在该模型上的实验验证了理论分析,并凸显了深度的有害影响。最后,我们通过跨编码器(Crosscoders)在实际模型中应用该框架,以顺序训练的Vision Transformer在CIFAR-10上的案例研究为例,展示了其分析能力。本工作为持续学习提供了以特征为中心的新视角。

原文摘要 · Abstract (English)

Catastrophic forgetting in continual learning is often measured at the performance or last-layer representation level, overlooking the underlying mechanisms. We introduce a mechanistic framework that offers a geometric interpretation of catastrophic forgetting as the result of transformations to the encoding of individual features. These transformations can lead to forgetting by reducing the allocated capacity of features or by disrupting their readout by downstream computations. Analysis of a tractable toy model formalizes this view, allowing us to identify best- and worst-case scenarios. Through experiments on this model, we empirically test our formal analysis and highlight the detrimental effect of depth. Finally, we demonstrate how our framework can be used in the analysis of practical models through the use of Crosscoders. We do so through a case study example of a Vision Transformer trained on sequential CIFAR-10. Our work provides a new, feature-centric vocabulary for continual learning.

持续学习可解释性特征编码视觉变压器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。