小模型在有限数据下,读经典书比儿童语料更有效。
What Should Baby Models Read? Exploring Sample-Efficient Data Composition on Model Performance
- 用1000万词数据集,测试不同语料对小模型的影响。
- 小模型在古腾堡语料上表现最好,儿童语料反而拖后腿。
- 模型大小决定最佳语料,需匹配数据与能力。
我们研究了预训练数据构成对小规模语言模型在样本高效设置下的影响。在仅1000万词的数据限制下,评估了儿童对话语料(CHILDES)、经典书籍(Gutenberg)、合成数据(TinyStories)及混合数据(Mix)在1800万至7.05亿参数模型上的表现。实验表明,较小模型(如GPT2-97M、GPT2-705M、Llama-360M)在更复杂丰富的语料如Gutenberg上表现更优;而基于CHILDES和TinyStories训练的模型在所有模型尺寸下均表现较差。结果表明,样本高效训练的最优数据集依赖于模型规模,且儿童语料与简化故事并非适用于所有规模的语言模型。强调了在高效训练中需同时考虑数据构成与模型容量。
原文摘要 · Abstract (English)
We explore the impact of pre-training data composition on the performance of small language models in a sample-efficient setting. Using datasets limited to 10 million words, we evaluate several dataset sources, including child-directed speech (CHILDES), classic books (Gutenberg), synthetic data (TinyStories), and a mix of these (Mix) across different model sizes ranging from 18 million to 705 million parameters. Our experiments show that smaller models (e.g., GPT2-97M, GPT2-705M, Llama-360M) perform better when trained on more complex and rich datasets like Gutenberg. Models trained on the CHILDES and TinyStories datasets underperformed across all model sizes. These findings suggest that the optimal dataset for sample efficient training depends on the model size, and that neither child-directed speech nor simplified stories are optimal for language models of all sizes. We highlight the importance of considering both dataset composition and model capacity for effective sample efficient language model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。