用高质量英文语料翻译构建德国语预训练数据集,提升德语模型性能。
KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report
- 将先进英文语料翻译为德语,保留文档边界与主题多样性。
- 在多个下游任务中,使用KletterMix训练的模型表现优于现有德语数据集。
- 适合德语NLP研究者、语言模型开发者及多语言研究者参考。
高质量预训练数据是现代语言模型的核心,但德语资源远不如英语丰富:规模小、整理不严、文档薄弱且缺乏受控实验验证。我们提出KletterMix,一个用于德语语言模型预训练与渐进训练的高质量语料库,作为可复用的数据资产供自然语言处理领域使用。KletterMix通过将最先进的英文预训练语料翻译为德语构建,同时保留文档边界、元数据、来源结构和主题多样性。该方法生成的德语文本在规模与多样性上达到现代预训练数据标准,并支持与原始英文语料的直接对比。我们通过广泛分析(包括翻译质量、文档长度分布、主题覆盖、来源构成与地理元数据)全面记录该数据集。利用COMETKiwi评估显示,译文在多个领域保持高质,表明精心翻译能有效保留原文语义与风格。此外,我们在控制实验中将KletterMix与既有德语语料对比,结果表明基于KletterMix训练的模型在德语下游任务中取得显著提升。这证明,经过精心策划的翻译数据能显著增强德语预训练数据生态。
原文摘要 · Abstract (English)
High-quality pretraining data is a central ingredient in modern language models, but German-language resources remain far less developed than their English counterparts: they are often smaller, less carefully curated, weakly documented, and rarely validated through controlled training experiments. We introduce KletterMix, a high-quality German corpus for language model pretraining and annealing, designed as a reusable dataset artifact for the natural language processing and modeling community. KletterMix is built by translating a state-of-the-art English pretraining corpus into German while preserving document boundaries, metadata, source structure, and topical diversity. This construction yields a German corpus with the scale and diversity of a modern pretraining dataset, while enabling direct comparison to its English source. We document the dataset through a broad set of corpus-level analyses, including translation quality, document length distributions, topic coverage, source composition, and geographic metadata. Using COMETKiwi, we show that the translated documents achieve strong quality across diverse domains, suggesting that careful translation can preserve much of the semantic and stylistic richness of the original corpus. Beyond dataset construction, we evaluate KletterMix as training data. Through controlled pretraining and annealing ablations against established German corpora, we show that models trained on KletterMix achieve measurable improvements on German-language downstream evaluations. These results demonstrate that carefully curated translated data can substantially strengthen the German pretraining data ecosystem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。