模型预训练中知识熵下降会阻碍新知识获取。
Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition
- 用知识熵衡量模型记忆来源的广度,熵越高表示越广泛利用记忆。
- 预训练越深入,知识熵越低,模型获取和保留新知识的能力下降。
- 激活沉寂记忆源可提升知识吸收能力,适合研究模型学习机制者阅读。
本文研究模型在预训练过程中参数化知识整合倾向的演变及其对性能的影响,尤其关注知识获取与遗忘问题。我们引入知识熵概念,量化模型所依赖的记忆来源范围:高熵表示广泛利用多种记忆源,低熵则意味着依赖少数确定性高的来源。分析显示,随着预训练推进,知识熵持续下降。这一下降与模型获取和保留知识能力减弱密切相关,表明知识熵降低(活跃记忆源减少)会损害知识获取与保持能力。进一步实验表明,增加沉寂记忆源的活跃度能提升模型的知识获取与保留能力。
原文摘要 · Abstract (English)
In this work, we investigate how a model's tendency to broadly integrate its parametric knowledge evolves throughout pretraining, and how this behavior affects overall performance, particularly in terms of knowledge acquisition and forgetting. We introduce the concept of knowledge entropy, which quantifies the range of memory sources the model engages with; high knowledge entropy indicates that the model utilizes a wide range of memory sources, while low knowledge entropy suggests reliance on specific sources with greater certainty. Our analysis reveals a consistent decline in knowledge entropy as pretraining advances. We also find that the decline is closely associated with a reduction in the model's ability to acquire and retain knowledge, leading us to conclude that diminishing knowledge entropy (smaller number of active memory sources) impairs the model's knowledge acquisition and retention capabilities. We find further support for this by demonstrating that increasing the activity of inactive memory sources enhances the model's capacity for knowledge acquisition and retention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。