打造5万亿词高质量语料库,助力小规模模型达到顶尖水平
Zyda-2: a 5 Trillion Token High-Quality Dataset
- 整合FineWeb、DCLM等优质开源数据,通过交叉去重与模型筛选提质
- 共含5万亿词,训练出的Zamba2系列模型在同类中表现领先
- 开源免费,适合研究者和开发者快速构建高效语言模型
本技术报告介绍Zyda-2:一个用于语言模型预训练的五万亿词数据集。Zyda-2由FineWeb和DCLM等高质量开源语料汇编而成,通过交叉去重和基于模型的质量过滤,提炼出最高质量子集。该数据集已用于训练Zamba2系列模型,其性能在同规模模型中处于领先地位。Zyda-2采用宽松开源许可,可在Hugging Face公开获取。
原文摘要 · Abstract (English)
In this technical report, we present Zyda-2: a five trillion token dataset for language model pretraining. Zyda-2 was used to train our Zamba2 series of models which are state-of-the-art for their weight class. We build Zyda-2 by collating high-quality open-source tokens such as FineWeb and DCLM, then distilling them to the highest-quality subset via cross-deduplication and model-based quality filtering. Zyda-2 is released under a permissive open license, and is available at https://huggingface.co/datasets/Zyphra/Zyda-2
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。