构建1540亿词元的开源德语文本库,助力开放模型训练
The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
- 整合41个来源的德语文本,覆盖七大领域
- 产出154.56亿词元高质量数据,支持模型训练
- 全数据集符合CC-BY-SA 4.0及以上许可,可复用
大语言模型训练依赖大规模语料,但多数数据授权不明确,尤其非英语语种更缺开源文本。本文提出德国共用语料库(German Commons),迄今最大的开源德语文本集合,涵盖法律、科学、文化、政治、新闻、经济与网络文本,来自41个来源的7个领域。通过系统化采集具有可验证授权的权威数据源,获得154.56亿词元的高质量文本。处理流程包含全面的质量过滤、去重与格式修复,确保异构数据的一致性。所有子集均采用至少CC-BY-SA 4.0或等效许可,保障模型训练与再分发的合规性。同时发布针对德语文本的语料构建与过滤代码,实现完全可复现与可扩展。该语料库填补了德语开源预训练数据的关键空白,推动真正开放的德语模型发展。
原文摘要 · Abstract (English)
Large language model development relies on large-scale training corpora, yet most contain data of unclear licensing status, limiting the development of truly open models. This problem is exacerbated for non-English languages, where openly licensed text remains critically scarce. We introduce the German Commons, the largest collection of openly licensed German text to date. It compiles data from 41 sources across seven domains, encompassing legal, scientific, cultural, political, news, economic, and web text. Through systematic sourcing from established data providers with verifiable licensing, it yields 154.56 billion tokens of high-quality text for language model training. Our processing pipeline implements comprehensive quality filtering, deduplication, and text formatting fixes, ensuring consistent quality across heterogeneous text sources. All domain subsets feature licenses of at least CC-BY-SA 4.0 or equivalent, ensuring legal compliance for model training and redistribution. The German Commons therefore addresses the critical gap in openly licensed German pretraining data, and enables the development of truly open German language models. We also release code for corpus construction and data filtering tailored to German language text, rendering the German Commons fully reproducible and extensible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。