提出数据质量新量化方法,可预测模型训练损失与算力需求
Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining
- 引入无量纲数据质量参数Q,统一建模数据量、模型规模与质量关系
- 实验显示高质量数据能显著降低模型规模和计算开销,效果可预测
- 提供两种实用估算方法,适合大模型训练中的数据筛选与资源配置
语言模型训练的缩放定律传统上描述性能随模型规模和数据量的变化。先前研究虽探讨了架构变体和数据处理(如数据过滤、噪声注入),但未在严谨的缩放定律框架内形式化数据质量。本文引入无量纲数据质量参数Q,提出一个考虑质量的扩展缩放定律,在Chinchilla框架基础上联合预测损失与模型大小、数据量及质量的关系。该定律基于有效样本量与信息论视角,解释噪声或冗余语料的影响,并提供两种实用估计器:(i) 噪声污染率代理,(ii) 缺失度量。通过神经机器翻译与自回归建模的合成实验,系统控制不同层次的噪声注入,验证了损失随数据质量可预测地变化;高质量数据可大幅减少所需模型规模与计算资源。结果表明有效数据量随质量呈次线性衰减,对中等程度数据污染具有鲁棒性。跨样本评估进一步证实该定律的预测能力。相比以往经验分析,本工作建立了一个明确且可推广的数据质量定律,为大规模预训练中数据精修投入与模型规模权衡提供了具体指导。
原文摘要 · Abstract (English)
Scaling laws for language model training traditionally characterize how performance scales with model size and dataset volume. Prior work has explored architecture variants and data treatments such as dataset filtering and noise injection in language model pretraining; however, these studies have not formalized data quality within a principled scaling law. We introduce a dimensionless data-quality parameter Q, and propose a quality-aware scaling law extending the Chinchilla framework to predict loss as a joint function of model size, data volume, and data quality. The law is motivated by an effective-sample-size and information-theoretic view of noisy or redundant corpora, and it admits two practical estimators for Q: (i) a corruption rate proxy and (ii) a deficiency measure. Through synthetic experiments in neural machine translation and autoregressive modeling -- where we systematically control data quality via multiple levels of noise injection variation -- we show that loss scales predictably with data quality and that higher-quality data can substantially reduce model size and hence compute requirements. Our results demonstrate a sublinear decay of effective data with quality and robustness to moderate data corruption; out-of-sample evaluations further validate the predictive form of the law. Unlike prior empirical analyses, our work establishes an explicit, generalizable law for data quality, offering concrete guidance for balancing data curation effort and model scale in large-scale pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。