用小学课程内容训练模型,研究知识如何被学习和使用。
LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

- 用880亿词的中小学教材构建可控训练数据集
- 50亿参数模型在该数据集上训练出可解释的能力边界
- 适合研究知识注入与能力扩展的实验场景
现代语言模型在异构网络文本上训练,导致难以刻画知识获取过程。为此,我们构建了包含880亿词的精选预训练语料库LITTLECURRICULUM,其内容严格限定于美国小学课程,排除五年级以上概念、事实和词汇。在此语料上从零训练出50亿参数的模型LITTLELEARNER,具备开放问答能力,但知识与能力范围清晰受限于可解释的课程指南。我们公开LITTLECURRICULUM和LITTLELEARNER,作为发展性受限的实验沙盒,用于研究模型在明确训练范围下的知识获取、表征与使用机制。通过首次实验验证,后训练和上下文学习可增强模型对已有知识的利用效率,但无法突破原有范围。研究结果凸显该受控环境对未来探究的价值。
原文摘要 · Abstract (English)
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LITTLECURRICULUM yields LITTLELEARNER, a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release LITTLECURRICULUM and LITTLELEARNER as a developmentally restricted sandbox to study how models acquire, represent, and use data under a well-defined training scope. We illustrate the sandbox's utility in a first suite of experiments on injecting new knowledge through post-training and in-context learning. These methods let LITTLELEARNER better utilize existing knowledge, but do not raise out-of-scope capabilities. Our findings underscore the value of this controlled environment for future investigations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。