arXiv:2603.12658cs.CLcs.AI2026-03被引 4

提出LLM持续学习框架,分三阶段解决知识遗忘问题。

Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages

  • 按预训练、微调、对齐三阶段设计持续学习框架
  • 对比分析传统方法在遗忘率与知识迁移上的表现
  • 适合关注大模型动态更新的研究者和开发者

持续学习(CL)已成为使大语言模型(LLMs)能够动态适应不断演化的知识和顺序任务的关键范式,同时缓解静态预训练带来的灾难性遗忘问题。本文系统梳理了针对LLMs的持续学习方法,围绕三个核心训练阶段:持续预训练、持续微调和持续对齐展开。在经典重放、正则化与架构方法基础上,进一步依据其遗忘缓解机制细分各类型,并对传统方法在适应性与关键改进方面进行严格对比分析。研究揭示了LLM持续学习与传统机器学习在规模、参数效率及涌现能力方面的核心差异。分析涵盖遗忘率、知识迁移效率等关键评估指标,以及新兴的性能评测基准。尽管现有方法表现良好,但实现无缝知识融合仍面临根本挑战,本文指出了关键开放问题与未来方向。

原文摘要 · Abstract (English)

Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting, a critical limitation of the static pre-training paradigm inherent to modern LLMs. This paper presents a comprehensive overview of CL methodologies tailored for LLMs, structured around three core training stages: continual pre-training, continual fine-tuning, and continual alignment. Beyond the canonical taxonomy of rehearsal-, regularization-, and architecture-based methods, we further subdivide each category by its distinct forgetting mitigation mechanisms and conduct a rigorous comparative analysis of the adaptability and critical improvements of traditional CL methods for LLMs. In doing so, we explicitly highlight core distinctions between LLM CL and traditional machine learning, particularly with respect to scale, parameter efficiency, and emergent capabilities. Our analysis covers essential evaluation metrics, including forgetting rates and knowledge transfer efficiency, along with emerging benchmarks for assessing CL performance. While current methods show promising results, fundamental challenges persist in achieving seamless knowledge integration, and we identify key open problems and future directions for the field.

持续学习大模型知识遗忘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。