arXiv:2411.06068cs.CLcs.AI2024-11被引 12

打造5万亿词高质量语料库,助力小规模模型达到顶尖水平

Zyda-2: a 5 Trillion Token High-Quality Dataset

  • 整合FineWeb、DCLM等优质开源数据,通过交叉去重与模型筛选提质
  • 共含5万亿词,训练出的Zamba2系列模型在同类中表现领先
  • 开源免费,适合研究者和开发者快速构建高效语言模型

本技术报告介绍Zyda-2:一个用于语言模型预训练的五万亿词数据集。Zyda-2由FineWeb和DCLM等高质量开源语料汇编而成,通过交叉去重和基于模型的质量过滤,提炼出最高质量子集。该数据集已用于训练Zamba2系列模型,其性能在同规模模型中处于领先地位。Zyda-2采用宽松开源许可,可在Hugging Face公开获取。

原文摘要 · Abstract (English)

In this technical report, we present Zyda-2: a five trillion token dataset for language model pretraining. Zyda-2 was used to train our Zamba2 series of models which are state-of-the-art for their weight class. We build Zyda-2 by collating high-quality open-source tokens such as FineWeb and DCLM, then distilling them to the highest-quality subset via cross-deduplication and model-based quality filtering. Zyda-2 is released under a permissive open license, and is available at https://huggingface.co/datasets/Zyphra/Zyda-2

数据集大模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。