arXiv:2502.17510cs.LGcs.AI2025-02ACL被引 14

动态识别参数重要性,缓解大模型持续学习中的遗忘问题。

Recurrent Knowledge Identification and Fusion for Language Model Continual Learning

  • 内环快速适应新任务,外环全局融合历史知识
  • 多轮迭代融合,显著减少灾难性遗忘
  • 适用于从0.77亿到130亿参数的各类语言模型

持续学习(CL)对于在动态现实环境中部署大语言模型(LLMs)至关重要,避免昂贵的重新训练。尽管基于参数重要性的模型集成与合并方法日益流行,但其常因依赖静态重要性估计而难以平衡知识迁移与遗忘。本文提出一种名为Recurrent-KIF的新框架,通过动态估计参数重要性分布,提升知识迁移能力。受人类持续学习启发,该框架采用内环快速适应新任务并识别关键参数,外环则通过冗余知识删减与关键知识融合,全局管理新旧知识的融合。内外环交替进行多轮融合,利用中间训练信息,并根据演化的重要程度自适应调整融合策略。在两个持续学习基准上,对多种模型规模(770M至13B)的大量实验表明,Recurrent-KIF有效缓解灾难性遗忘,增强知识迁移能力。

原文摘要 · Abstract (English)

Continual learning (CL) is crucial for deploying large language models (LLMs) in dynamic real-world environments without costly retraining. While recent model ensemble and model merging methods guided by parameter importance have gained popularity, they often struggle to balance knowledge transfer and forgetting, mainly due to the reliance on static importance estimates during sequential training. In this paper, we present Recurrent-KIF, a novel CL framework for Recurrent Knowledge Identification and Fusion, which enables dynamic estimation of parameter importance distributions to enhance knowledge transfer. Inspired by human continual learning, Recurrent-KIF employs an inner loop that rapidly adapts to new tasks while identifying important parameters, coupled with an outer loop that globally manages the fusion of new and historical knowledge through redundant knowledge pruning and key knowledge merging. These inner-outer loops iteratively perform multiple rounds of fusion, allowing Recurrent-KIF to leverage intermediate training information and adaptively adjust fusion strategies based on evolving importance distributions. Extensive experiments on two CL benchmarks with various model sizes (from 770M to 13B) demonstrate that Recurrent-KIF effectively mitigates catastrophic forgetting and enhances knowledge transfer.

持续学习大模型知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。