arXiv:2607.22662cs.AI2026-07

CuraWeb通过联合优化质量、冗余与多样性,构建更全面的网页预训练数据集。

CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data

论文配图:CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data
图 1 · 摘自论文原文
  • 采用双轨清洗与混合去重,协同优化数据质量与分布多样性。
  • 在30亿参数模型上,平均性能提升1.8%,尤其增强长尾知识与推理能力。
  • 适合追求高质量、广覆盖预训练数据的研究者与工业应用者。

通过高度筛选过滤构建的开放网页语料库(如FineWeb-Edu和DCLM)是大语言模型预训练数据的核心,显著提升了模型性能。然而,这些流程通常依赖单一优化目标,不可避免地缩小了数据分布多样性,弱化了长尾知识,限制了数据覆盖范围并浪费了开放网络的巨大潜力。为此,我们提出一种新的数据筛选范式,从线性删减转向质量、冗余与多样性的联合优化。该框架融合双轨清洗(规则驱动与模型驱动)与混合去重(n-gram与语义),并采用多目标采样器平衡信息质量与分布广度。基于Common Crawl构建的CuraWeb是一个2万亿标记的英文语料库。与现有资源不同,CuraWeb以工业级标准恢复更完整的数据分布,显著提升多样性且冗余极低,实现对多样化领域中长尾知识的更广泛覆盖。30亿参数规模的实验表明,CuraWeb在多种基准测试中显著优于现有基线,平均性能提升1.8%,尤其在知识密集型与推理任务中表现突出。

原文摘要 · Abstract (English)

Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly advanced LLM performance. However, these pipelines typically rely on singular optimization objectives, which inevitably narrows distributional diversity and marginalizes long-tail knowledge, thereby restricting data coverage and underutilizing the vast potential of the open web. To address this limitation, we propose a novel curation paradigm that shifts from linear pruning to the joint optimization of quality, redundancy, and diversity. This framework synergizes dual-track cleaning (rule-based and model-driven) with hybrid deduplication (n-gram and semantic), while employing a multi-objective sampler to balance informational quality with distributional breadth. Applying this framework to Common Crawl, we construct CuraWeb, a 2T-token English corpus. Unlike existing resources, CuraWeb establishes an industrial-grade standard for data curation by recovering a more holistic data distribution with enhanced diversity and minimal redundancy, achieving broader coverage of long-tail knowledge across diverse domains. Experimental evaluations at the 3B scale demonstrate that CuraWeb significantly outperforms state-of-the-art baselines, yielding an average performance gain of 1.8\% across a wide range of benchmarks, particularly in knowledge-intensive and reasoning tasks.

数据清洗预训练多样性长尾知识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。