首个开放许可的荷兰语大模型训练数据集,含360亿独有荷兰语标记
GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training
- 整合21个荷兰语专属数据源,共36B未被其他大模型使用过的荷兰语标记
- 包含约2070亿英文、2320亿代码、480亿德语/丹麦语标记,经合规筛选
- 所有数据开源可商用,适合开发合法、无害的荷兰语语言模型
我们提出GPT-NL公共语料库,是目前最大的开放许可荷兰语资源集合。该语料库包含21个纯荷兰语数据集,总计360亿经预处理的荷兰语标记,且未出现在任何其他大模型预训练语料中。此外,还包含约2070亿英文、2320亿代码、480亿德语/丹麦语标记,源自现有数据集并经过合规性筛选。数据涵盖Common Crawl、Common Corpus等大型语料库,以及与机构合作采集或合成生成的全新荷兰语内容。所有数据均遵循宽松许可,经评估后以CC-BY协议重新发布。完整数据集已公开在Hugging Face Hub。
原文摘要 · Abstract (English)
We present the GPT-NL Public Corpus, the biggest permissively licensed corpus of Dutch language resources. The GPT-NL Public Corpus contains 21 Dutch-only collections totalling 36B preprocessed Dutch tokens not present in any other LLM pretraining corpus. Additionally, the corpus includes roughly 207B English, 232B Code, and 48B German/Danish tokens taken from existing sets which we further curated for compliance. This corpus includes curated data from large existing corpora like Common Corpus and Common Crawl, as well as newly created Dutch-specific collections. Most newly created Dutch collections consist of content collected in collaboration with organisations or synthetically augmented content. All data is collected and evaluated with the aim of facilitating the creation of (commercial) language models that are lawful, useful and non-harmful. All data included in the GPT-NL Public Corpus is sourced from datasets with permissive licensing and is curated and redistributed under a CC-BY license. The full dataset is publicly available on the Hugging Face Hub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。