arXiv:2605.16374cs.LGcs.AI2026-05

揭示视觉模型中概念遗忘的本质是可访问性下降而非信息消失

Lost or Hidden? A Concept-Level Forgetting in Supervised Continual Learning

论文配图:Lost or Hidden? A Concept-Level Forgetting in Supervised Continual Learning
图 1 · 摘自论文原文
  • 用稀疏自编码器构建任务锚定的潜在特征空间,将每个隐变量视为视觉概念代理
  • 发现多数看似丢失的概念信息在假设线性条件下可恢复,且随着任务增加解码能力下降
  • 为持续学习提供新视角:遗忘主因是表示可访问性降低,适合研究模型内部机制者阅读

持续学习研究模型如何在不丢失旧知识的前提下适应新任务。尽管已有大量方法缓解灾难性遗忘,但该领域仍以性能为导向,缺乏对视觉模型表征空间中遗忘本质的深入理解。现有研究多通过任务级性能或粗粒度表征漂移分析遗忘,未能区分输出可访问性与内部结构变化。为此,本文提出一种诊断框架,利用稀疏自编码器(SAEs)构建任务锚定的潜在特征空间,将单个SAE隐变量视为模型内部计算中反复出现且相对解耦的视觉模式的语义代理。在此框架下,我们将遗忘分解为表观概念删除、可恢复性和可解码性。结果显示,大部分看似丢失的概念级信息在假设线性条件下通常可恢复,且随着任务数量增加,概念可解码性逐渐下降。总体表明,相当一部分概念级遗忘源于表征可访问性的改变,而非信息完全擦除。

原文摘要 · Abstract (English)

Continual learning studies how models can adapt to new tasks while retaining previously acquired knowledge. Although a broad spectrum of methods has been proposed to mitigate catastrophic forgetting, the field remains predominantly performance-driven, with limited insight into what forgetting actually corresponds to within the vision model's representation space. Prior work has primarily analyzed forgetting through task-level performance or coarse measures of representational drift, without disentangling output-level accessibility from changes in finer-grained internal structure. To this end, we propose a diagnostic framework that leverages Sparse Autoencoders (SAEs) to define a task-anchored latent feature space, enabling analysis of how task-specific information evolves at a finer granularity, where individual SAE latents are treated as concept proxies for recurring and relatively disentangled visual patterns in the model's internal computations. Within this framework, we decompose forgetting into apparent concept deletion, recoverability, and decodability. We show that a large portion of seemingly lost concept-level information can often be recovered under linearity assumption, with concept decodability degrading as more tasks are introduced. Overall, our findings suggest that a significant part of concept-level forgetting can be attributed to changes in the representational accessibility rather than complete information erasure.

持续学习表征分析概念遗忘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。