arXiv:2412.15285cs.CLcs.AI2024-12被引 23

通过两阶段预训练提升大模型准确率,实测效果优于随机排序17%。

Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining

  • 分两阶段设计数据混合策略,先小规模实验再放大到15万亿token
  • 相比自然数据分布,平均准确率提升17%,最大提升达3.4%
  • 适合想优化训练数据组合的模型开发者和科研人员

有效预训练大语言模型依赖于数据选择、混合与排序的策略。然而,由于模型开发者披露有限,数据混合方案在更长文本跨度和更大模型规模下的可扩展性仍不明确。为此,本文正式提出两阶段预训练概念,并系统研究如何配置数据以最大化模型准确率。结果表明,两阶段方法相较随机数据排序和自然令牌分布,平均准确率分别提升3.4%和17%。我们提供基于数据源质量与训练轮次的数据混合设计指南,并验证了在1万亿令牌规模下设计的混合方案可有效扩展至15万亿令牌及250亿参数模型。这些发现为从业者设计和扩展数据混合方案提供了可操作路径。

原文摘要 · Abstract (English)

Pretraining large language models effectively requires strategic data selection, blending and ordering. However, key details about data mixtures especially their scalability to longer token horizons and larger model sizes remain underexplored due to limited disclosure by model developers. To address this, we formalize the concept of two-phase pretraining and conduct an extensive systematic study on how to select and mix data to maximize model accuracies for the two phases. Our findings illustrate that a two-phase approach for pretraining outperforms random data ordering and natural distribution of tokens by 3.4% and 17% on average accuracies. We provide in-depth guidance on crafting optimal blends based on quality of the data source and the number of epochs to be seen. We propose to design blends using downsampled data at a smaller scale of 1T tokens and then demonstrate effective scaling of our approach to larger token horizon of 15T tokens and larger model size of 25B model size. These insights provide a series of steps practitioners can follow to design and scale their data blends.

大模型训练数据混合两阶段预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。