揭示大模型如何通过知识电路演化吸收新知识
How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training
- 从知识电路视角分析持续预训练中的计算子图演变
- 发现新知识获取受已有知识相关性影响,存在形成到优化的阶段跃迁
- 知识电路演化呈现由深层向浅层发展的规律,适合研究模型内部机制者参考
尽管大型语言模型在知识密集型任务中表现优异,但对其如何内化新知识、尤其是如何在神经计算中结构化存储知识的理解仍显不足。本文从知识电路演化的视角出发,识别出促进知识存储与处理的计算子图。通过对持续预训练过程中电路演化的系统分析,发现:(1)新知识的获取受其与已有知识的相关性影响;(2)知识电路演化经历从形成到优化的显著阶段跃迁;(3)知识电路演化遵循由深层到浅层的发展模式。这些发现不仅深化了对大模型新知识获取机制的理论认知,也为改进持续预训练策略以提升模型性能提供了潜在启示。代码与数据将公开于 https://github.com/zjunlp/DynamicKnowledgeCircuits。
原文摘要 · Abstract (English)
Despite exceptional capabilities in knowledge-intensive tasks, Large Language Models (LLMs) face a critical gap in understanding how they internalize new knowledge, particularly how to structurally embed acquired knowledge in their neural computations. We address this issue through the lens of knowledge circuit evolution, identifying computational subgraphs that facilitate knowledge storage and processing. Our systematic analysis of circuit evolution throughout continual pre-training reveals several key findings: (1) the acquisition of new knowledge is influenced by its relevance to pre-existing knowledge; (2) the evolution of knowledge circuits exhibits a distinct phase shift from formation to optimization; (3) the evolution of knowledge circuits follows a deep-to-shallow pattern. These insights not only advance our theoretical understanding of the mechanisms of new knowledge acquisition in LLMs, but also provide potential implications for improving continual pre-training strategies to enhance model performance. Code and data will be available at https://github.com/zjunlp/DynamicKnowledgeCircuits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。