arXiv:2506.16884cs.LGcs.AI2025-06ICML被引 7

模型越懒,持续学习效果越好;适度特征学习才能避免遗忘。

The Importance of Being Lazy: Scaling Limits of Continual Learning

  • 用可调参数化区分‘懒’与‘富’训练模式,揭示规模影响机制。
  • 宽度增大仅在减少特征学习时有效,能显著降低灾难性遗忘。
  • 发现任务相似性决定模型从‘懒’到‘富’的临界点,适合长期学习研究者。

尽管已有诸多努力,神经网络在非平稳环境中仍难以持续学习,对灾难性遗忘(CF)的理解尚不完整。本文系统研究了模型规模和特征学习程度对持续学习的影响。通过可变参数化架构,区分‘懒’与‘富’训练范式,化解文献中关于规模效应的矛盾观点。结果表明,增加模型宽度仅在降低特征学习程度、提升‘懒度’时才有效。基于动力学平均场理论,我们分析了无限宽度下特征学习状态的动态行为,刻画了灾难性遗忘,拓展了以往局限于‘懒’状态的理论结果。研究揭示:高特征学习仅在任务高度相似时有益;当任务非平稳性增强,存在由任务相似性调控的相变,模型从低遗忘的‘懒’态跃迁至高遗忘的‘富’态。最终发现,模型在依赖任务非平稳性的临界特征学习水平上表现最优,且该最优水平可跨模型规模迁移。本工作为规模与特征学习在持续学习中的作用提供了统一视角。

原文摘要 · Abstract (English)

Despite recent efforts, neural networks still struggle to learn in non-stationary environments, and our understanding of catastrophic forgetting (CF) is far from complete. In this work, we perform a systematic study on the impact of model scale and the degree of feature learning in continual learning. We reconcile existing contradictory observations on scale in the literature, by differentiating between lazy and rich training regimes through a variable parameterization of the architecture. We show that increasing model width is only beneficial when it reduces the amount of feature learning, yielding more laziness. Using the framework of dynamical mean field theory, we then study the infinite width dynamics of the model in the feature learning regime and characterize CF, extending prior theoretical results limited to the lazy regime. We study the intricate relationship between feature learning, task non-stationarity, and forgetting, finding that high feature learning is only beneficial with highly similar tasks. We identify a transition modulated by task similarity where the model exits an effectively lazy regime with low forgetting to enter a rich regime with significant forgetting. Finally, our findings reveal that neural networks achieve optimal performance at a critical level of feature learning, which depends on task non-stationarity and transfers across model scales. This work provides a unified perspective on the role of scale and feature learning in continual learning.

持续学习灾难性遗忘模型规模特征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。