用有效训练量衡量数据质量,提升小模型性能
Scaling Parameter-Constrained Language Models with Quality Data
- 提出有效训练标记数概念,融合文本多样性与合成度
- 模型在200个规模下验证,预测准确率相关性达+0.83
- 适合关注小模型优化与数据质量评估的研究者
语言建模中的缩放定律传统上将训练损失与数据规模和模型参数关联,提供计算最优估计,但常忽略数据质量对模型泛化的影响。本文在原有框架中引入微观视角,提出有效训练标记数作为关键性能指标,其由文本多样性与教师模型测得的合成度两个可计算指标组合而成。我们在多样化的采样合成数据上预训练了超过200个参数量介于2500万至15亿之间的模型,并估算了文本质量、模型规模、训练标记数与八项推理任务准确率之间的常数关系。结果表明,估算常数与真实准确率间的皮尔逊相关系数达+0.83,且在常用数据采样与合成等提升数据质量的技术场景中具有分析价值。
原文摘要 · Abstract (English)
Scaling laws in language modeling traditionally quantify training loss as a function of dataset size and model parameters, providing compute-optimal estimates but often neglecting the impact of data quality on model generalization. In this paper, we extend the conventional understanding of scaling law by offering a microscopic view of data quality within the original formulation -- effective training tokens -- which we posit to be a critical determinant of performance for parameter-constrained language models. Specifically, we formulate the proposed term of effective training tokens to be a combination of two readily-computed indicators of text: (i) text diversity and (ii) syntheticity as measured by a teacher model. We pretrained over $200$ models of 25M to 1.5B parameters on a diverse set of sampled, synthetic data, and estimated the constants that relate text quality, model size, training tokens, and eight reasoning task accuracy scores. We demonstrated the estimated constants yield +0.83 Pearson correlation with true accuracies, and analyzed it in scenarios involving widely-used data techniques such as data sampling and synthesis which aim to improve data quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。