arXiv:2502.10442cs.LGcs.AI2025-02被引 3

过参数化可缓解持续学习中的灾难性遗忘,理论证明有效。

Analysis of Overparameterization in Continual Learning under a Linear Model

  • 在线性回归模型中,仅靠过参数化即可减少遗忘。
  • 当参数量远超数据量时,首个任务的预测风险显著降低。
  • 适用于研究持续学习与双下降现象的理论分析者。

能够按顺序学习多个任务的自主机器学习系统易出现灾难性遗忘问题。为理解持续学习中的遗忘程度,亟需数学理论支持。作为迈向这一目标的基础步骤,本文从理论视角研究了无显式防遗忘机制的梯度下降下的持续学习。在受启发于置换任务的两任务设置中,我们通过解析推导证明:当过参数化比例足够高时,顺序训练两个任务的模型对第一个任务仍能获得低风险估计器。本工作还建立了单个线性回归任务的风险非渐近界,可能对双下降理论领域具有独立价值。

原文摘要 · Abstract (English)

Autonomous machine learning systems that learn many tasks in sequence are prone to the catastrophic forgetting problem. Mathematical theory is needed in order to understand the extent of forgetting during continual learning. As a foundational step towards this goal, we study continual learning and catastrophic forgetting from a theoretical perspective in the simple setting of gradient descent with no explicit algorithmic mechanism to prevent forgetting. In this setting, we analytically demonstrate that overparameterization alone can mitigate forgetting in the context of a linear regression model. We consider a two-task setting motivated by permutation tasks, and show that as the overparameterization ratio becomes sufficiently high, a model trained on both tasks in sequence results in a low-risk estimator for the first task. As part of this work, we establish a non-asymptotic bound of the risk of a single linear regression task, which may be of independent interest to the field of double descent theory.

持续学习过参数化线性模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。