提出参数更新幅度理论模型,解释遗忘与泛化关系
Exploring the Impact of Parameter Update Magnitude on Forgetting and Generalization of Continual Learning
- 从参数空间漂移角度建模遗忘机制,揭示更新幅度影响
- 理论推导出最小遗忘的最优更新幅度,统一两类训练范式
- 设计自适应更新策略,在多个数据集上超越传统方法
参数更新幅度被认为是持续学习中的关键因素。然而,现有研究多集中于设计多样化的更新策略,对内在机制的理论理解仍显不足。本文从参数更新幅度的角度刻画模型遗忘,将其形式化为任务特定参数空间漂移导致的知识退化,这一现象此前因假设统一参数空间而未被充分捕捉。通过推导最小化遗忘的最优更新幅度,本文将冻结训练与初始化训练两种典型范式统一到受限参数更新的优化框架中。理论分析进一步表明,在参数距离较小的任务序列下,冻结训练相比初始化训练展现出更优的泛化能力与更低的遗忘率。这些理论洞见启发了一种基于梯度方向自适应调整更新幅度的新混合策略。在深度神经网络上的实验表明,该方法显著优于标准训练策略,为设计高效可扩展的持续学习算法提供了新的理论视角与实践指导。
原文摘要 · Abstract (English)
The magnitude of parameter updates are considered a key factor in continual learning. However, most existing studies focus on designing diverse update strategies, while a theoretical understanding of the underlying mechanisms remains limited. Therefore, we characterize model's forgetting from the perspective of parameter update magnitude and formalize it as knowledge degradation induced by task-specific drift in the parameter space, which has not been fully captured in previous studies due to their assumption of a unified parameter space. By deriving the optimal parameter update magnitude that minimizes forgetting, we unify two representative update paradigms, frozen training and initialized training, within an optimization framework for constrained parameter updates. Our theoretical results further reveals that sequence tasks with small parameter distances exhibit better generalization and less forgetting under frozen training rather than initialized training. These theoretical insights inspire a novel hybrid parameter update strategy that adaptively adjusts update magnitude based on gradient directions. Experiments on deep neural networks demonstrate that this hybrid approach outperforms standard training strategies, providing new theoretical perspectives and practical inspiration for designing efficient and scalable continual learning algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。