用模型筛选+生成数据,打造6280亿词德语训练集
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation
- 结合启发式与模型筛选,生成3290亿词合成数据
- 在10亿和80亿参数模型上均超越FineWeb2
- 适合德语大模型研究者与工业级应用开发者
扩大数据量对大语言模型至关重要,但近期研究显示数据质量能显著提升性能与训练效率。本文提出一种德语数据集构建流程,融合启发式与模型驱动的过滤技术,并生成合成数据。基于该流程构建了包含6280亿词的Aleph-Alpha-GermanWeb数据集,由三部分组成:(1) Common Crawl网络数据(780亿词),(2) FineWeb2有机数据(2350亿词),(3) 基于真实网页数据生成的合成数据(3290亿词)。我们通过从头预训练10亿参数的Llama风格模型和80亿参数的无分词器层级自回归变换器(HAT)验证该数据集。在包括MMMLU在内的德语基准测试中,Aleph-Alpha-GermanWeb相比仅使用FineWeb2有显著性能提升,即使在将FineWeb2与维基百科等人工精选高质量数据源结合后,该优势仍保持在80亿参数规模。结果支持模型驱动的数据清洗与合成可有效增强大模型预训练数据。
原文摘要 · Abstract (English)
Scaling data quantity is essential for large language models (LLMs), yet recent findings show that data quality can significantly boost performance and training efficiency. We introduce a German-language dataset curation pipeline that combines heuristic and model-based filtering techniques with synthetic data generation. We use our pipeline to create Aleph-Alpha-GermanWeb, a 628B-word German pre-training dataset composed of three subsets drawing from: (1) Common Crawl web data (organic subset; 78B words), (2) FineWeb2 (organic subset; 235B), and (3) synthetically-generated data conditioned on actual, organic web data (synthetic subset; 329B). We evaluate our dataset by pre-training both a 1B Llama-style model and an 8B tokeniser-free hierarchical autoregressive transformer (HAT) from scratch. A comparison on German-language benchmarks, including MMMLU, shows significant performance gains of Aleph-Alpha-GermanWeb over FineWeb2 alone. This advantage holds at the 8B scale even when FineWeb2 is enriched by human-curated high-quality data sources such as Wikipedia. Our findings support the growing body of evidence that model-based data curation and synthetic data generation can significantly enhance LLM pre-training datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。