研究神经网络在任务持续变化时的内核动态,发现静态内核假设在持续学习中不成立。
Reactivation: Empirical NTK Dynamics Under Task Shifts
- 通过实证分析任务切换下的NTK演化,揭示持续学习中的动态特性。
- 发现即使在大规模模型下,静态内核近似仍无法准确描述训练过程。
- 为持续学习提供了新的理论检验视角,适合关注模型动态的学者。
神经正切核(NTK)是研究神经网络功能动态的强大工具。在所谓的懒惰或核区,训练过程中NTK保持静态,网络函数在静态神经正切特征空间中呈线性。而NTK的演化是特征学习的关键,也是深度学习成功的重要驱动力。近年来对NTK动态的研究推动了泛化与缩放行为的重要发现。然而,这些工作主要局限于单一任务设定,即数据分布随时间保持不变。本文首次对持续学习中数据分布随时间变化时的NTK动态进行了全面的实证分析。研究结果表明,持续学习是一个丰富且未被充分挖掘的测试平台,可用于深入探究神经网络训练的动态机制。同时,研究挑战了在持续学习理论分析中使用静态内核近似的有效性,即使在大规模模型下亦然。
原文摘要 · Abstract (English)
The Neural Tangent Kernel (NTK) offers a powerful tool to study the functional dynamics of neural networks. In the so-called lazy, or kernel regime, the NTK remains static during training and the network function is linear in the static neural tangents feature space. The evolution of the NTK during training is necessary for feature learning, a key driver of deep learning success. The study of the NTK dynamics has led to several critical discoveries in recent years, in generalization and scaling behaviours. However, this body of work has been limited to the single task setting, where the data distribution is assumed constant over time. In this work, we present a comprehensive empirical analysis of NTK dynamics in continual learning, where the data distribution shifts over time. Our findings highlight continual learning as a rich and underutilized testbed for probing the dynamics of neural training. At the same time, they challenge the validity of static-kernel approximations in theoretical treatments of continual learning, even at large scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。