arXiv:2601.03570cs.CL2026-01被引 3

揭示大模型在持续预训练中概念学习与遗忘的内在机制

How Do Large Language Models Learn Concepts During Continual Pre-Training?

  • 通过概念电路分析模型内部表征动态
  • 发现概念学习呈先升后稳的阶段性模式
  • 提出电路感知的重放策略以缓解遗忘

人类主要通过概念(如狗)理解世界,这些抽象心智表征结构化了感知、推理与学习。然而,大语言模型(LLMs)在持续预训练过程中如何获取、保留与遗忘这些概念仍不清晰。本文研究单个概念的习得与遗忘过程,以及多个概念间的干扰与协同作用。我们关联行为动态与模型内部的概念电路——与特定概念相关的计算子图,并引入图论指标刻画电路拓扑。分析发现:(1) 概念电路提供非平凡且一致的学与忘信号;(2) 概念电路呈现阶段性时间模式,先上升后渐降并趋于稳定;(3) 学习增益大的概念在后续训练中更易遗忘;(4) 语义相似概念间干扰更强;(5) 概念知识转移能力不同,部分显著促进其他概念学习。研究为概念学习动态提供了电路层面的视角,启发了如电路感知经验重放等概念感知训练策略。

原文摘要 · Abstract (English)

Human beings primarily understand the world through concepts (e.g., dog), abstract mental representations that structure perception, reasoning, and learning. However, how large language models (LLMs) acquire, retain, and forget such concepts during continual pretraining remains poorly understood. In this work, we study how individual concepts are acquired and forgotten, as well as how multiple concepts interact through interference and synergy. We link these behavioral dynamics to LLMs' internal concept circuits, computational subgraphs associated with specific concepts, and incorporate graph metrics to characterize circuit topology. Our analysis reveals: (1) LLMs concept circuits provide a non-trivial, consistent signal of concept learning and forgetting; (2) concept circuits exhibit a stage-wise temporal pattern during continual pretraining, with an early increase followed by gradual decrease and stabilization; (3) concepts with larger learning gains tend to exhibit greater forgetting under subsequent training; (4) semantically similar concepts induce stronger interference than weakly related ones; (5) conceptual knowledge differs in their transferability, with some significantly facilitating the learning of others. Together, our findings provide a circuit-level view of concept learning dynamics and motivate concept-aware training strategies, such as Circuit-aware Experience Replay, which uses circuit topology to prioritize concepts vulnerable to forgetting.

概念学习大模型持续学习电路分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。