首次从线性变换视角分析持续学习中的表征遗忘问题。
Measuring Representational Shifts in Continual Learning: A Linear Transformation Perspective
- 提出表征差异度量,量化模型不同训练阶段的隐藏层表征变化。
- 发现高层网络遗忘更快,网络宽度增加可缓解遗忘现象。
- 适用于研究持续学习机制及模型设计的科研人员。
在持续学习场景中,先前任务的灾难性遗忘是关键挑战,因此有效衡量此类遗忘至关重要。近年来,研究者愈发关注表征遗忘——即在隐藏层层面测量的遗忘。本文首次对表征遗忘进行理论分析,并基于该分析深入理解持续学习行为。首先,我们引入一种新度量——表征差异,用于衡量通过持续学习训练的模型在两个时间点的表征空间差异。我们证明该度量能有效替代表征遗忘,且具备良好的可解析性。其次,通过对该度量的数学分析,我们得出若干关键发现:遗忘程度随网络层数增加而加剧,而增加网络宽度可减缓遗忘过程。最后,我们在真实图像数据集(包括 Split-CIFAR100 与 ImageNet1K)上验证了这些理论结果。
原文摘要 · Abstract (English)
In continual learning scenarios, catastrophic forgetting of previously learned tasks is a critical issue, making it essential to effectively measure such forgetting. Recently, there has been growing interest in focusing on representation forgetting, the forgetting measured at the hidden layer. In this paper, we provide the first theoretical analysis of representation forgetting and use this analysis to better understand the behavior of continual learning. First, we introduce a new metric called representation discrepancy, which measures the difference between representation spaces constructed by two snapshots of a model trained through continual learning. We demonstrate that our proposed metric serves as an effective surrogate for the representation forgetting while remaining analytically tractable. Second, through mathematical analysis of our metric, we derive several key findings about the dynamics of representation forgetting: the forgetting occurs more rapidly to a higher degree as the layer index increases, while increasing the width of the network slows down the forgetting process. Third, we support our theoretical findings through experiments on real image datasets, including Split-CIFAR100 and ImageNet1K.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。