发现深度模型持续学习失效源于新任务初始化时曲率坍缩,提出双正则化方法缓解。
Spectral Collapse Drives Loss of Plasticity in Deep Continual Learning
- 通过分析海森矩阵谱坍缩,揭示梯度下降失效机制。
- 在多个持续学习任务中,结合两种正则化使模型保持可塑性。
- 适合研究持续学习、神经网络优化与正则化方向的读者。
我们研究深度神经网络在持续学习中丧失可塑性的原因,发现其失败前会出现新任务初始化时的海森矩阵谱坍缩,即有意义的曲率方向消失,导致梯度下降失效。通过对线性化ReLU网络的分析,推导出成功训练的ε-秩条件,并证明损失加权格拉姆矩阵与广义高斯-牛顿近似谱等价,从而将NTK动态与海森曲率关联。针对谱坍缩,我们引入海森矩阵的克罗内克积近似,提出两种正则化改进:维持高有效特征秩和应用L2惩罚。在持续监督学习与强化学习任务上的实验表明,结合这两种正则化能有效保留模型可塑性。
原文摘要 · Abstract (English)
We investigate why deep neural networks suffer from loss of plasticity in continual learning, and thus fail to learn new tasks without reinitializing parameters. We show that this failure is preceded by Hessian spectral collapse at new-task initialization, where meaningful curvature directions vanish and gradient descent becomes ineffective. Analyzing a linearized ReLU network, we derive explicit $ε$-rank conditions for successful training and prove that the loss-weighted Gram matrix is spectrally equivalent to the Generalized Gauss-Newton approximation, thereby relating NTK dynamics to Hessian curvature. Targeting spectral collapse directly, we then discuss the Kronecker factored approximation of the Hessian, which motivates two regularization enhancements: maintaining high effective feature rank and applying L2 penalties. Experiments on continual supervised and reinforcement learning tasks confirm that combining these two regularizers effectively preserves plasticity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。