arXiv:2410.18505cs.CL2024-10被引 14

构建500GB高质量中文语料库,助力大模型预训练

CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models

论文配图:CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models
图 1 · 摘自论文原文
  • 采用两阶段混合过滤流程提升数据质量
  • 用100B词元训练0.5B参数模型,零样本性能超越多个基准
  • 适合中文大模型研究者和开发者使用

我们提出CCI3.0-HQ(https://huggingface.co/datasets/BAAI/CCI3-HQ),是基于中文网络语料库CCI3.0(https://huggingface.co/datasets/BAAI/CCI3-Data)的500GB高质量子集,通过新颖的两阶段混合过滤管道显著提升数据质量。为评估其有效性,我们在100B词元的数据上从头训练了一个0.5B参数模型,在10个基准测试中实现零样本优于CCI3.0、SkyPile和WanjuanV1的表现。高质量过滤过程有效将Qwen2-72B-instruct模型的能力提炼至0.5B模型,获得中文网页数据分类的最优F1分数。该开源数据集有望促进高质量语言模型的广泛可及性。

原文摘要 · Abstract (English)

We present CCI3.0-HQ (https://huggingface.co/datasets/BAAI/CCI3-HQ), a high-quality 500GB subset of the Chinese Corpora Internet 3.0 (CCI3.0)(https://huggingface.co/datasets/BAAI/CCI3-Data), developed using a novel two-stage hybrid filtering pipeline that significantly enhances data quality. To evaluate its effectiveness, we trained a 0.5B parameter model from scratch on 100B tokens across various datasets, achieving superior performance on 10 benchmarks in a zero-shot setting compared to CCI3.0, SkyPile, and WanjuanV1. The high-quality filtering process effectively distills the capabilities of the Qwen2-72B-instruct model into a compact 0.5B model, attaining optimal F1 scores for Chinese web data classification. We believe this open-access dataset will facilitate broader access to high-quality language models.

中文语料大模型预训练数据清洗开源数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。