词频重叠能预测模型在主流评测上的表现。
Benchmarks Are Not That Out of Distribution: Word Overlap Predicts Performance
- 用词级单字熵衡量训练与测试数据的词汇重叠度。
- 重叠度越高,零样本评测得分越低,且效果随数据量提升。
- 适合关注评测可信度与数据分布影响的研究者。
理解高质量预训练数据的构成仍是语言模型训练的核心问题。本文研究了基准评测性能是否主要由预训练语料与评估数据之间的统计模式重叠程度决定。我们通过词级单字熵和词频统计来衡量这种重叠,并在10个零样本基准、4个涵盖85亿至600亿词元的预训练数据集以及4亿到30亿参数的模型规模下进行受控实验。结果表明,词级单字熵与基准性能之间存在稳健的反向关系,说明广泛使用的基准评测受训练与测试数据间词汇重叠的显著影响。更大的预训练子集若具有相似的词级单字熵,则能带来更好的下游性能,表明词频统计在塑造评测分数方面也起着额外作用。综合来看,许多标准基准与预训练语料的分布差异较弱,因此简单的词汇重叠统计即可有效预测评测表现。
原文摘要 · Abstract (English)
Understanding what constitutes high-quality pre-training data remains a central question in language model training. In this work, we investigate whether benchmark performance is primarily driven by the degree of statistical pattern overlap between pre-training corpora and evaluation datasets. We measure this overlap using word-level unigram cross-entropy and word frequency statistics, and perform controlled experiments across $10$ zero-shot benchmarks, $4$ pre-training datasets spanning $8.5\mathrm{B}$ to $60\mathrm{B}$ tokens, and model sizes ranging from $400\mathrm{M}$ to $3\mathrm{B}$ parameters. Our results demonstrate a robust inverse relationship between word-level unigram cross-entropy and benchmark performance, suggesting that widely used benchmarks are strongly influenced by word overlap between training and evaluation data. Thus, larger pre-training subsets with similar word-level unigram cross-entropy yield improved downstream results, indicating that word frequency statistics play an additional role in shaping benchmark scores. Taken together, these results suggest that many standard benchmarks are only weakly out-of-distribution relative to pre-training corpora, so that simple word-overlap statistics predict benchmark performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。