构建高质量中英双语数据集,提升大模型推理能力
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models
- 通过多阶段去重与质量评分,清洗35TB中英混合数据
- 提取45亿条思维链模板,减少幻觉并覆盖多样推理路径
- 特别适合需要强逻辑推理的数学与代码任务研究者
我们提出CCI4.0,一个大规模双语预训练数据集,旨在提升大语言模型的类人推理能力。该数据集约占用35TB磁盘空间,包含两个子集:CCI4.0-M2-Base和CCI4.0-M2-CoT。前者整合了5.2TB经人工精筛的中文网络语料、22.5TB来自Nemotron-CC的英文数据,以及数学、维基、arXiv和代码等多样化来源。尽管数据主要来自已处理数据集,各领域质量标准动态变化,需大量专家经验与人力处理。为此,我们设计了一种基于模型的新流水线,通过两阶段去重、多分类器质量评分和领域感知流畅性过滤来保障数据质量。我们从中提取出45亿条思维链(CoT)模板,命名为CCI4.0-M2-CoT。与从大模型蒸馏CoT不同,本方法分阶段提取,更充分展现多样化的推理模式,显著降低幻觉风险。实证评估表明,使用CCI4.0预训练的大模型获得更清洁、可靠的训练信号,在下游任务中表现更优,尤其在数学与代码反思任务上提升明显。结果凸显严格数据清洗与人类思维模板对提升大模型性能的关键作用,为自动处理预训练语料提供了新思路。
原文摘要 · Abstract (English)
We introduce CCI4.0, a large-scale bilingual pre-training dataset engineered for superior data quality and diverse human-like reasoning trajectory. CCI4.0 occupies roughly $35$ TB of disk space and comprises two sub-datasets: CCI4.0-M2-Base and CCI4.0-M2-CoT. CCI4.0-M2-Base combines a $5.2$ TB carefully curated Chinese web corpus, a $22.5$ TB English subset from Nemotron-CC, and diverse sources from math, wiki, arxiv, and code. Although these data are mostly sourced from well-processed datasets, the quality standards of various domains are dynamic and require extensive expert experience and labor to process. So, we propose a novel pipeline justifying data quality mainly based on models through two-stage deduplication, multiclassifier quality scoring, and domain-aware fluency filtering. We extract $4.5$ billion pieces of CoT(Chain-of-Thought) templates, named CCI4.0-M2-CoT. Differing from the distillation of CoT from larger models, our proposed staged CoT extraction exemplifies diverse reasoning patterns and significantly decreases the possibility of hallucination. Empirical evaluations demonstrate that LLMs pre-trained in CCI4.0 benefit from cleaner, more reliable training signals, yielding consistent improvements in downstream tasks, especially in math and code reflection tasks. Our results underscore the critical role of rigorous data curation and human thinking templates in advancing LLM performance, shedding some light on automatically processing pretraining corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。