arXiv:2606.18024cs.LGcs.AI2026-06中稿 · ICML被引 1

揭示持续学习中遗忘的低秩本质,提出可精准预测遗忘方向的新理论。

Catastrophic Forgetting is Low-Rank: A Function-Space Theory for Continual Adaptation

  • 从函数空间出发,用核方法建模新旧任务间的遗忘机制。
  • 遗忘集中在少数旧任务特征模式上,且在特定结构下呈可预测的低秩特性。
  • 适用于模型微调、参数高效训练等场景,为防御遗忘提供新思路。

持续适应中的灾难性遗忘通常通过参数漂移、回放或知识蒸馏来研究,但这些方法无法识别哪些输出空间方向易受干扰。本文在核极限(NTK)框架下给出函数空间解释:新任务训练通过跨任务核导致旧任务预测漂移,从而在任意新任务梯度更新前,闭式推导出遗忘向量。在固定主干线性头的参数高效微调(PEFT-CL)中,该预测精确至数值精度;对于非线性适配器或全微调,则为局部NTK近似。相同表达式表明,遗忘集中在少量旧任务NTK特征模式上,在固定线性头时还导出脆弱秩的克罗内克缩放规则。这些结果澄清了与先前NTK重叠理论的关系,解释了为何参数空间正则化可能忽略输出空间干扰,并启发了一种针对性的谱正则化方法。

原文摘要 · Abstract (English)

Catastrophic forgetting in continual adaptation is usually studied through parameter drift, replay, or distillation, but these views do not identify which output-space directions are vulnerable. We give a function-space account in the NTK regime: new-task training induces old-task prediction drift through the cross-task kernel, yielding a closed-form predictor for the forgetting vector before any new-task gradient step. In frozen-backbone linear-head PEFT-CL, where the model is linear in the trainable parameters, the predictor is exact up to numerical precision; for nonlinear adapters/full fine-tuning, it is a local NTK approximation. The same expression reveals that forgetting concentrates in a small number of old-task NTK eigenmodes and under frozen linear heads gives a Kronecker scaling rule for the vulnerable rank. These results clarify the relation to prior NTK-overlap theory, explain why parameter-space regularizers can miss output-space interference, and motivate a targeted spectral regularizer.

持续学习遗忘抑制函数空间低秩结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。