arXiv:2504.04039cs.LGcs.AI2025-04被引 4

通过结构正则化实现记忆与性能的权衡,缓解持续学习中的灾难性遗忘。

Memory-Statistics Tradeoff in Continual Learning with Structural Regularization

  • 基于前任务海森矩阵设计自适应正则化,减少遗忘。
  • 正则化向量越多,记忆开销越大,但误差越小。
  • 性能接近联合训练,适合需要高效增量学习的场景。

我们研究了在理想随机设计下两个线性回归任务的持续学习问题的统计性能。采用一种结构正则化算法,利用针对前一任务海森矩阵的广义ℓ₂正则化来缓解灾难性遗忘。本文建立了该算法联合过失误差的上下界。分析揭示了记忆复杂度与统计效率之间的根本权衡:正则化中向量数量增加会恶化记忆复杂度,但降低过失误差;反之亦然。此外,理论表明无正则化的朴素持续学习会引发灾难性遗忘,而结构正则化能有效缓解此问题。值得注意的是,该方法性能可媲美同时访问两个任务的联合训练。这些结果突显了曲率感知正则化在持续学习中的关键作用。

原文摘要 · Abstract (English)

We study the statistical performance of a continual learning problem with two linear regression tasks in a well-specified random design setting. We consider a structural regularization algorithm that incorporates a generalized $\ell_2$-regularization tailored to the Hessian of the previous task for mitigating catastrophic forgetting. We establish upper and lower bounds on the joint excess risk for this algorithm. Our analysis reveals a fundamental trade-off between memory complexity and statistical efficiency, where memory complexity is measured by the number of vectors needed to define the structural regularization. Specifically, increasing the number of vectors in structural regularization leads to a worse memory complexity but an improved excess risk, and vice versa. Furthermore, our theory suggests that naive continual learning without regularization suffers from catastrophic forgetting, while structural regularization mitigates this issue. Notably, structural regularization achieves comparable performance to joint training with access to both tasks simultaneously. These results highlight the critical role of curvature-aware regularization for continual learning.

持续学习正则化统计分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。