arXiv:2604.13460cs.LGcs.AI2026-04被引 1

从任务顺序转向分布,揭示持续学习中遗忘的谱结构机制

From Order to Distribution: A Spectral Characterization of Forgetting in Continual Learning

论文配图:From Order to Distribution: A Spectral Characterization of Forgetting in Continual Learning
图 1 · 摘自论文原文
  • 将任务视为分布采样,而非固定顺序,研究遗忘的理论规律
  • 推导出遗忘量的精确算子恒等式,揭示其递归谱结构
  • 发现遗忘速度由任务分布几何特性决定,适用于理论分析者

持续学习中的核心挑战是遗忘——在顺序学习新任务时对旧任务性能的下降。尽管遗忘已被广泛实证研究,但严谨的理论刻画仍有限。近期工作分析了在过参数线性回归中,固定任务集合在随机顺序下的遗忘行为。本文转换视角,从顺序转为分布:不考虑固定任务集合在随机顺序下的表现,而是研究任务从任务分布Π独立同分布采样下的精确拟合线性情形,探究生成分布本身如何影响遗忘。在此设定下,我们推导出遗忘量的精确算子恒等式,揭示其递归谱结构。基于该恒等式,我们建立了无条件上界,识别出主导渐近项,并在一般非退化情形下刻画了收敛速率(至常数)。进一步地,我们将该速率与任务分布的几何性质关联,阐明该模型中遗忘快慢的驱动因素。

原文摘要 · Abstract (English)

A central challenge in continual learning is forgetting, the loss of performance on previously learned tasks induced by sequential adaptation to new ones. While forgetting has been extensively studied empirically, rigorous theoretical characterizations remain limited. A notable step in this direction is \citet{evron2022catastrophic}, which analyzes forgetting under random orderings of a fixed task collection in overparameterized linear regression. We shift the perspective from order to distribution. Rather than asking how a fixed task collection behaves under random orderings, we study an exact-fit linear regime in which tasks are sampled i.i.d.\ from a task distribution~$Π$, and ask how the generating distribution itself governs forgetting. In this setting, we derive an exact operator identity for the forgetting quantity, revealing a recursive spectral structure. Building on this identity, we establish an unconditional upper bound, identify the leading asymptotic term, and, in generic nondegenerate cases, characterize the convergence rate up to constants. We further relate this rate to geometric properties of the task distribution, clarifying what drives slow or fast forgetting in this model.

持续学习遗忘机制谱分析理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。