arXiv:2502.14907cs.CLcs.AI2025-02被引 7

GneissWeb构建了10万亿词的高质量数据集,助力大模型训练性能提升。

GneissWeb: Preparing High Quality Data for LLMs at Scale

  • 采用分片精确子串去重与多级质量过滤器组合策略
  • 训练模型在11项基准测试中领先细调Web数据集2.73个百分点
  • 适合追求高精度大模型训练的研究者与工业应用

数据量与质量对大语言模型(LLM)性能至关重要。高质量数据能显著提升模型在各类下游任务上的泛化能力。当前主流大模型的预训练数据集尚未公开,而多数开源数据集规模较小(不足5万亿词),难以支撑大模型训练。本文提出GneissWeb,一个约含10万亿词的数据集,满足大模型对数据数量与质量的要求。其构建方法包括分片精确子串去重和精心设计的多级质量过滤器组合。GneissWeb在数据质量与数量间取得良好平衡,所训练模型在11个常用基准测试(包含零样本与少样本)上平均得分优于现有开源最大数据集(5+万亿词)2.73个百分点;当扩展至20个基准测试时,仍领先1.75个百分点。

原文摘要 · Abstract (English)

Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's ability to generalize on a wide range of downstream tasks. Large pre-training datasets for leading LLMs remain inaccessible to the public, whereas many open datasets are small in size (less than 5 trillion tokens), limiting their suitability for training large models. In this paper, we introduce GneissWeb, a large dataset yielding around 10 trillion tokens that caters to the data quality and quantity requirements of training LLMs. Our GneissWeb recipe that produced the dataset consists of sharded exact sub-string deduplication and a judiciously constructed ensemble of quality filters. GneissWeb achieves a favorable trade-off between data quality and quantity, producing models that outperform models trained on state-of-the-art open large datasets (5+ trillion tokens). We show that models trained using GneissWeb dataset outperform those trained on FineWeb-V1.1.0 by 2.73 percentage points in terms of average score computed on a set of 11 commonly used benchmarks (both zero-shot and few-shot) for pre-training dataset evaluation. When the evaluation set is extended to 20 benchmarks (both zero-shot and few-shot), models trained using GneissWeb still achieve a 1.75 percentage points advantage over those trained on FineWeb-V1.1.0.

数据集构建大模型训练语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。