arXiv:2502.10341cs.CL2025-02ICML被引 79

把网页按主题和格式分类,让预训练数据更有效。

Organize the Web: Constructing Domains Enhances Pre-Training Data Curation

  • 用主题与格式构建网页分类体系,实现自动标注。
  • 不同领域数据混合能显著提升下游任务性能。
  • 适合关注数据质量与结构化的模型训练研究者。

现代语言模型使用包含数万亿标记的大型非结构化网络数据集进行训练。由于缺乏结构,难以系统分析内容并优化数据筛选。本文通过构建内容分类体系,将原始网络语料拆解为多个领域,提出WebOrganizer框架,基于主题与格式对网页进行组织。利用大语言模型提炼出高效分类器,自动标注预训练数据。我们研究了不同领域数据混合策略对下游任务的影响,发现结合有效主题与格式可进一步提升模型表现。实验表明,该方法还能增强基于质量的数据选择方法的效果。此外,我们分析了质量筛选方法如何隐式改变领域分布。结果表明,领域构建与混合是质量筛选的重要补充,为更有效、可解释的预训练数据优化开辟新路径。

原文摘要 · Abstract (English)

Modern language models are trained on large, unstructured datasets consisting of trillions of tokens and obtained by crawling the web. The unstructured nature makes it difficult to reason about their contents and develop systematic approaches to data curation. In this paper, we unpack monolithic web corpora by developing taxonomies of their contents and organizing them into domains. We introduce WebOrganizer, a framework for organizing web pages in terms of both their topic and format. Using these two complementary notions of domains, we automatically annotate pre-training data by distilling annotations from a large language model into efficient classifiers. This allows us to study how data from different domains should be mixed to improve models on downstream tasks, and we show that we can combine insights about effective topics and formats to further boost performance. We demonstrate that our domain mixing also improves existing methods that select data based on quality. Furthermore, we study and compare how quality-based methods will implicitly change the domain mixture. Overall, our work demonstrates that constructing and mixing domains provides a valuable complement to quality-based data curation methods, opening new avenues for effective and insightful pre-training data curation.

数据清洗预训练领域划分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。