提出新指标优化准备度,更准预测模型持续学习能力。
Predicting Plasticity in Deep Continual Learning: A Theoretical Perspective

- 用梯度强度与可靠性构建新指标‘优化准备度’
- 实验证明该指标比旧方法更可靠预测训练潜力
- 适合研究持续学习泛化性与模型诊断的学者
深度持续学习要求模型在不从头训练的情况下适应新任务。然而,神经网络在学习完先前任务后可能丧失适应新任务的能力,即“可塑性损失”。现有多种解释和诊断方法被提出。受“所有模型皆错,但有些有用”启发,我们探究现有诊断能否预测模型可塑性。本文从实用角度将可塑性理解为可训练性,即模型在未来任务上的优化增益。理论上,通过构造反例,证明了广泛使用的表示秩和神经正切核秩等诊断方法在回归与分类场景中均可能失效。为此,我们提出新指标——优化准备度,结合梯度强度与梯度可靠性。在标准光滑性假设下,证明其可下界一步优化增益,提供理论保证。实验表明,在慢变回归与置换MNIST等常见设置中,优化准备度能更可靠地按可训练性排序检查点,且样本量显著更少。
原文摘要 · Abstract (English)
Deep continual learning requires models to adapt to new tasks without retraining from scratch. However, neural networks can lose their ability to adapt to new tasks after training on previous ones, a phenomenon known as loss of plasticity. There have been several explanations and diagnostics proposed for plasticity loss. Motivated by the philosophy that "all models are wrong, but some are useful", we ask: can existing diagnostics predict a neural network's plasticity? In this work, we take a practical view to interpret plasticity as trainability, i.e., a neural network's future optimization gain on a target task. We first take a theoretical approach, showing, by constructing a few counterexamples, that some widely adopted diagnostics of plasticity, including representation rank and neural tangent kernel rank, can fail to predict the loss of trainability in both regression and classification settings. We instead propose a novel metric, called optimization readiness, which combines gradient strength and gradient reliability. We prove that optimization readiness lower bounds one-step optimization gain under standard smoothness assumptions, providing a theoretical guarantee for its predictive power. Empirically, we show that across commonly used deep continual learning settings, such as Slowly-Changing Regression and Permuted MNIST, optimization readiness more reliably ranks checkpoints by trainability than prior diagnostics, even with substantially fewer samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。