揭示大模型学知识的三阶段过程及幻觉成因
How do language models learn facts? Dynamics, curricula and hallucinations
- 发现模型学知识分三阶段,先形成注意力回路再精准记忆
- 数据分布不均会缩短学习平台期,影响知识积累效率
- 微调易破坏已有知识,适合研究知识稳定性的学者
大型语言模型在预训练中积累了海量知识,但其获取机制尚不明确。本文通过合成事实回忆任务研究模型学习动态,发现三个关键结果:第一,模型学习分为三个阶段,在获得精确知识前经历性能平台期;机制上,该平台期与支持回忆的注意力回路形成同步。第二,训练数据分布显著影响学习动态,不平衡分布导致平台期缩短。第三,幻觉与知识同时出现,通过微调引入新知识时极易破坏原有参数化记忆。研究强调数据分布对知识获取的重要性,提出新型数据调度策略以加速神经网络训练。
原文摘要 · Abstract (English)
Large language models accumulate vast knowledge during pre-training, yet the dynamics governing this acquisition remain poorly understood. This work investigates the learning dynamics of language models on a synthetic factual recall task, uncovering three key findings: First, language models learn in three phases, exhibiting a performance plateau before acquiring precise factual knowledge. Mechanistically, this plateau coincides with the formation of attention-based circuits that support recall. Second, the training data distribution significantly impacts learning dynamics, as imbalanced distributions lead to shorter plateaus. Finally, hallucinations emerge simultaneously with knowledge, and integrating new knowledge into the model through fine-tuning is challenging, as it quickly corrupts its existing parametric memories. Our results emphasize the importance of data distribution in knowledge acquisition and suggest novel data scheduling strategies to accelerate neural network training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。