arXiv:2506.20920cs.CL2025-06被引 146

自动适配千种语言的预训练数据清洗流水线,提升多语种大模型性能。

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

  • 基于FineWeb设计可自动适配任意语言的数据清洗流程。
  • 在9种语言上验证,生成模型性能优于现有非英语数据集。
  • 支持超1000语言,产出20TB、50亿文档的多语种数据集。

训练顶尖大语言模型需海量清洁且多样化的文本数据。尽管高质量英文预训练数据的开放已取得显著进展,但训练高性能多语种大模型仍面临挑战,主要源于为众多语言定制过滤与去重流水线的复杂性。本文提出一种基于FineWeb的新数据集构建流程,可自动适配任意语言。我们在九种不同语言上系统评估了流水线设计选择,依据通过可衡量标准筛选出的有意义评估任务。结果表明,该流程能生成性能优于以往非英语数据集的语料库。我们还提出一种兼顾重复度与质量的简洁均衡方法,进一步提升性能。最终,将该流程扩展至1000余种语言,利用近100个Common Crawl快照构建了FineWeb2——一个20TB(50亿文档)的多语言数据集,并同步开源流水线、训练与评估代码。

原文摘要 · Abstract (English)

Pre-training state-of-the-art large language models (LLMs) requires vast amounts of clean and diverse text data. While the open development of large high-quality English pre-training datasets has seen substantial recent progress, training performant multilingual LLMs remains a challenge, in large part due to the inherent difficulty of tailoring filtering and deduplication pipelines to a large number of languages. In this work, we introduce a new pre-training dataset curation pipeline based on FineWeb that can be automatically adapted to support any language. We extensively ablate our pipeline design choices on a set of nine diverse languages, guided by a set of meaningful and informative evaluation tasks that were chosen through a novel selection process based on measurable criteria. Ultimately, we show that our pipeline can be used to create non-English corpora that produce more performant models than prior datasets. We additionally introduce a straightforward and principled approach to rebalance datasets that takes into consideration both duplication count and quality, providing an additional performance uplift. Finally, we scale our pipeline to over 1000 languages using almost 100 Common Crawl snapshots to produce FineWeb2, a new 20 terabyte (5 billion document) multilingual dataset which we release along with our pipeline, training, and evaluation codebases.

多语言数据清洗大模型预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。