arXiv:2606.08452cs.LG2026-06中稿 · Transactions on Ma…被引 1

用控制理论设计新方法,让模型持续学习不遗忘。

Theoretical Foundations of Continual Learning via Drift-Plus-Penalty

论文配图:Theoretical Foundations of Continual Learning via Drift-Plus-Penalty
图 1 · 摘自论文原文
  • 基于随机优化的漂移-惩罚原理,动态调节学习过程。
  • 在多个基准上表现优于现有方法,遗忘可控且可调。
  • 适合需要长期稳定学习的场景,如在线系统迭代。

现实世界中数据流具有非平稳性且按序到达,要求学习系统持续适应而无需从头训练。持续学习(CL)通过融入新任务同时缓解灾难性遗忘(即学习新知识导致旧知识性能下降)来应对这一挑战。本文提出一种控制论视角下的持续学习框架,显式调控遗忘演化过程,将适应视为受长期稳定性约束的受控过程。聚焦于基于回放的持续学习,即有限记忆缓冲区存储先前任务的代表性样本。提出基于漂移-惩罚(DPP)原则的连续学习方法COLD,同时引入作为参考基准的COLD-ORACLE变体。在每个任务中,两种方法均最小化当前任务损失,并维护一个虚拟队列,追踪对以往任务的长期稳定性偏差,从而将稳定与可塑性的权衡建模为受控的动力学过程。我们建立了稳定性与收敛性保证,通过可调控制参数刻画该权衡。实验表明,COLD在标准基准上持续优于多种先进方法,同时通过显式调节稳定性与可塑性,实现竞争性且可控的遗忘行为。

原文摘要 · Abstract (English)

In many real-world settings, data streams are nonstationary and arrive sequentially, requiring learning systems to adapt continuously without retraining from scratch. Continual learning (CL) addresses this challenge by incorporating new tasks while mitigating catastrophic forgetting, where learning new information degrades performance on previously acquired knowledge. We introduce a control-theoretic perspective on CL that explicitly regulates the evolution of forgetting, framing adaptation as a controlled process subject to long-term stability constraints. We focus on replay-based CL, where a finite memory buffer stores representative samples from prior tasks. We propose COntinual Learning with Drift-Plus-Penalty (COLD), a continual learning framework based on the Drift-Plus-Penalty (DPP) principle from stochastic optimization. To facilitate analysis, we also consider an oracle variant, COLD-ORACLE, as a reference benchmark. At each task, both methods minimize the current task loss while maintaining a virtual queue that tracks deviations from long-term stability on previously learned tasks, capturing the stability-plasticity trade-off as a regulated dynamical process. We establish stability and convergence guarantees that characterize this trade-off through a tunable control parameter. Experiments on standard benchmarks demonstrate that COLD consistently outperforms a broad range of state-of-the-art CL methods while providing competitive and controllable forgetting behavior through explicit regulation of stability and plasticity.

持续学习控制理论遗忘控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。