首次给出神经网络持续学习遗忘的闭式理论保障。
On the Theory of Continual Learning with Gradient Descent for Neural Networks
- 用梯度下降分析单隐层二次网络在正交聚类任务中的学习动态。
- 揭示训练中遗忘速率受迭代次数、样本量、任务数和隐藏层宽度共同影响。
- 理论结合实验,适用于研究模型稳定性与长期学习机制的研究者。
持续学习指模型在不遗忘旧知识的前提下适应连续任务序列,是人工智能的核心目标。为深入理解其内在机制,本文在可解析且具代表性的情况下分析持续学习的局限性:考虑用梯度下降训练单隐层二次神经网络,在带有高斯噪声的XOR-聚类数据集序列上学习,不同任务对应均值正交的聚类。通过紧致刻画梯度下降训练损失的动力学行为,推导出训练期遗忘速率关于迭代次数、样本量、任务数量和隐藏层宽度的显式上界。进一步利用算法稳定性框架,对泛化误差差距进行约束,从而获得测试期遗忘的相应保证。结果首次为神经网络持续学习中的遗忘现象提供了闭式理论保障,并揭示关键问题参数如何共同调控遗忘动态。数值实验验证了理论预测。
原文摘要 · Abstract (English)
Continual learning, the ability of a model to adapt to an ongoing sequence of tasks without forgetting earlier ones, is a central goal of artificial intelligence. To better understand its underlying mechanisms, we study the limitations of continual learning in a tractable yet representative setting. Specifically, we analyze one-hidden-layer quadratic neural networks trained by gradient descent on a sequence of XOR-cluster datasets with Gaussian noise, where different tasks correspond to clusters with orthogonal means. Our analysis is based on a tight characterization of gradient descent dynamics for the training loss, which yields explicit bounds on the rate of train-time forgetting as functions of the number of iterations, sample size, number of tasks, and hidden-layer width. We then leverage an algorithmic stability framework to bound the generalization gap, leading to corresponding guarantees on test-time forgetting. Together, our results provide the first closed-form guarantees for forgetting in continual learning with neural networks and show how key problem parameters jointly govern forgetting dynamics. Numerical experiments corroborate our theoretical results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。