arXiv:2601.07636cs.LG2026-01AAAI

提出新方法分解梯度尖锐性,提升持续学习泛化能力且计算轻量。

Beyond Sharpness: A Flatness Decomposition Framework for Efficient Continual Learning

  • 将尖锐性扰动拆分为梯度对齐与噪声成分,仅保留噪声项促进泛化
  • 在有限训练时间内仍保持显著性能优势,相比传统优化器提升10%以上
  • 可无缝集成到多种持续学习框架,适合资源受限场景部署

持续学习旨在让模型在不遗忘旧知识的前提下顺序学习多个任务。近期研究表明,优化至更平坦的损失极小值能提升模型泛化能力。然而,现有针对持续学习的尖锐性感知方法存在两大局限:(1) 将尖锐性正则视为单一信号,未区分其组成部分的贡献;(2) 引入显著计算开销,阻碍实际部署。为此,我们提出FLAD——一种新型优化框架,将尖锐性感知扰动分解为梯度对齐与随机噪声成分,并证明仅保留噪声成分即可促进泛化。我们进一步设计轻量级调度策略,使FLAD在训练时间受限时仍能保持显著性能提升。该方法可无缝集成于各类持续学习范式,在多样实验设置中一致优于标准及尖锐性感知优化器,验证了其有效性与实用性。

原文摘要 · Abstract (English)

Continual Learning (CL) aims to enable models to sequentially learn multiple tasks without forgetting previous knowledge. Recent studies have shown that optimizing towards flatter loss minima can improve model generalization. However, existing sharpness-aware methods for CL suffer from two key limitations: (1) they treat sharpness regularization as a unified signal without distinguishing the contributions of its components. and (2) they introduce substantial computational overhead that impedes practical deployment. To address these challenges, we propose FLAD, a novel optimization framework that decomposes sharpness-aware perturbations into gradient-aligned and stochastic-noise components, and show that retaining only the noise component promotes generalization. We further introduce a lightweight scheduling scheme that enables FLAD to maintain significant performance gains even under constrained training time. FLAD can be seamlessly integrated into various CL paradigms and consistently outperforms standard and sharpness-aware optimizers in diverse experimental settings, demonstrating its effectiveness and practicality in CL.

持续学习优化方法泛化能力轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。