arXiv:2410.00868cs.LG2024-10

通过灵活约束梯度方向,显著提升持续学习中记忆旧知识与学习新知识的平衡。

Fine-Grained Gradient Restriction: A Simple Approach for Mitigating Catastrophic Forgetting

  • 提出更灵活的梯度限制方法,取代固定强度的记忆约束。
  • 在多个数据集上实现更优的旧知识保留与新任务学习权衡。
  • 计算高效,适合实际场景中的持续学习应用。

持续学习的核心挑战在于平衡学习新任务与保留旧知识之间的关系。梯度情景记忆(GEM)通过利用部分历史样本约束模型参数更新方向来实现这一平衡。本文分析了GEM中常被忽视的超参数——记忆强度,发现其能提升性能主要因其增强了GEM的泛化能力,从而带来更优的权衡效果。基于此,我们提出了两种更灵活的更新方向约束方法,相较使用记忆强度,在所有测试数据集上均实现了更优的帕累托前沿(Pareto Frontier),即在记忆与学习之间取得更好平衡。此外,我们还设计了一种计算高效的近似求解方法,以应对更多约束下的优化问题。

原文摘要 · Abstract (English)

A fundamental challenge in continual learning is to balance the trade-off between learning new tasks and remembering the previously acquired knowledge. Gradient Episodic Memory (GEM) achieves this balance by utilizing a subset of past training samples to restrict the update direction of the model parameters. In this work, we start by analyzing an often overlooked hyper-parameter in GEM, the memory strength, which boosts the empirical performance by further constraining the update direction. We show that memory strength is effective mainly because it improves GEM's generalization ability and therefore leads to a more favorable trade-off. By this finding, we propose two approaches that more flexibly constrain the update direction. Our methods are able to achieve uniformly better Pareto Frontiers of remembering old and learning new knowledge than using memory strength. We further propose a computationally efficient method to approximately solve the optimization problem with more constraints.

持续学习梯度约束记忆保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。