对德语数据重复使用高质量过滤数据,比一次遍历大量低质数据更高效。
Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling

- 构建分层质量过滤器,从5亿网页文档中提取高信号数据
- 重复训练高质量数据集,在多模型规模下性能持续领先7轮
- 适合追求高效训练的非英语大模型研究者参考
近期研究表明,对大规模英文网络语料进行高质量子集筛选可显著提升训练效率。然而,对于德语、法语或日语等高资源非英语语言,激进过滤带来策略困境:是优先选择多样性,用大量轻度过滤数据单次训练,还是优先保证质量,严格筛选高质量核心数据并多次重复训练?本文针对德语构建了应用于5亿网页文档的分层质量过滤体系,比较了在过滤子集上多轮训练与在广泛语料上单次训练的效果。实验覆盖多个模型规模和词元预算,结果表明:重复高质量数据始终优于单次训练更大、较宽松过滤的数据集。值得注意的是,即使经过7轮训练,性能差距依然存在。研究显示,对非英语大模型而言,通过质量过滤实现语义集中,比单纯扩大数据量更利于高效语言建模。我们发布了名为Boldt的德语语言模型及清理后的评估基准,实验表明其在训练仅使用10-360倍更少词元的情况下,仍达到顶尖水平。
原文摘要 · Abstract (English)
Recent research has shown that filtering massive English web corpora into high-quality subsets significantly improves training efficiency. However, for high-resource non-English languages like German, French, or Japanese, aggressive filtering creates a strategic dilemma: should practitioners prioritize diversity by training once on large amounts of lightly filtered web data, or prioritize quality by strictly filtering for a high-quality core and repeating it over multiple epochs? We investigate this trade-off for German by constructing hierarchical quality filters applied to 500M web documents, comparing multi-epoch training on the filtered subsets against single-pass training on a diverse corpus. Our experiments across multiple model scales and token budgets show that repeating high-quality data consistently outperforms single-pass training on larger, less filtered sets. Notably, the performance gap persists even after 7 epochs. Our findings suggest that for non-English LLMs, semantic concentration through quality filtering offers a more viable path to efficient language modeling than simply maximizing unique data volume. We release our German language models (called Boldt), as well as our cleaned evaluation benchmarks to the research community. Our experiments indicate that they achieve state-of-the-art results despite training on 10-360x fewer tokens than comparable models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。