arXiv:2609.04180cs.CLcs.AI2026-09

大模型预训练中,多样化知识表达能提升学习效果。

Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views

论文配图:Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views
图 1 · 摘自论文原文
  • 用改写文本作为辅助视角,帮助模型更好吸收知识。
  • 固定词元预算下,分配更多资源给改写文本可提升事实记忆能力。
  • 即使教师模型较弱,改写文本仍有效,适合研究预训练机制的人。

目前对大语言模型在预训练中如何获取知识的理解仍不完整。我们提出,辅助视角(即知识的重述形式)对学习具有因果促进作用,并设计控制实验加以验证。首先,确认重复是知识获取的必要条件,且改写仅在小批量时有帮助。其次,在固定词元预算下,将文档重复的资源转向辅助视角,能提升学习效果,甚至改善事实回忆。第三,辅助视角的有效性不依赖于生成它们的教师模型强度。第四,我们识别出两类知识形式——上下文性和基础性知识——在存在知识空白时有助于学习。最后,通过分层偏差与压缩机制分析其内在机理。结果表明,自然出现在大规模预训练语料中的知识辅助表示,是预训练成功的关键因素之一,也为数据多样性的重要性提供了合理解释。

原文摘要 · Abstract (English)

Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning, counterintuitively, even for factual recall. Third, the effectiveness of auxiliary views is not contingent on the strength of the teacher model that generates them. Fourth, we identify forms of knowledge, contextual and foundational, that aid learning in the presence of prior knowledge gaps. Finally, we examine how these effects manifest mechanistically via layer-wise biases and compression. Together, our findings suggest that auxiliary representations of knowledge, which arise naturally in large pre-training corpora, are a key factor in the success of pre-training and offer a plausible explanation for why data diversity matters.

预训练知识获取大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。