提出自适应模型融合方法,提升持续学习的稳定性与泛化能力。
BECAME: BayEsian Continual Learning with Adaptive Model MErging
- 基于贝叶斯原理重构融合机制,自动计算最优融合系数。
- 在多个基准数据集上超越现有持续学习方法,显著减少遗忘。
- 适合需要长期增量学习的场景,如智能助手、机器人视觉。
持续学习(CL)旨在跨任务增量学习的同时缓解灾难性遗忘。其核心挑战在于平衡稳定性(保留旧知识)与可塑性(学习新任务)。现有梯度投影方法虽能保证稳定性,但常抑制可塑性;而模型融合技术虽具潜力,却依赖经验假设和人工调参。本文探索模型融合在优化稳定性-可塑性权衡中的作用,提供理论支撑,并基于贝叶斯持续学习框架重构融合机制,推导出随任务特性自适应的最优融合系数闭式解。为验证该方法,提出两阶段框架BECAME,融合梯度投影与自适应融合的优势。大量实验表明,该方法优于当前主流持续学习方法及现有融合策略。
原文摘要 · Abstract (English)
Continual Learning (CL) strives to learn incrementally across tasks while mitigating catastrophic forgetting. A key challenge in CL is balancing stability (retaining prior knowledge) and plasticity (learning new tasks). While representative gradient projection methods ensure stability, they often limit plasticity. Model merging techniques offer promising solutions, but prior methods typically rely on empirical assumptions and carefully selected hyperparameters. In this paper, we explore the potential of model merging to enhance the stability-plasticity trade-off, providing theoretical insights that underscore its benefits. Specifically, we reformulate the merging mechanism using Bayesian continual learning principles and derive a closed-form solution for the optimal merging coefficient that adapts to the diverse characteristics of tasks. To validate our approach, we introduce a two-stage framework named BECAME, which synergizes the expertise of gradient projection and adaptive merging. Extensive experiments show that our approach outperforms state-of-the-art CL methods and existing merging strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。