arXiv:2506.01732cs.CL2025-06被引 19

构建了全球最大开源语言模型训练数据集,含两万亿无版权文本。

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

  • 收集无版权或开放许可数据,总规模达两万亿标记符。
  • 覆盖多语言及代码数据,支持多语种与跨领域预训练。
  • 适合关注合规性、开源研究与多语言模型开发的团队使用。

大型语言模型(LLMs)通常在来自不同来源和领域的海量数据上进行预训练,这些数据常包含数万亿个标记符,其中大量内容受版权或专有权限保护,引发其法律使用的争议。这凸显了对真正开放且符合数据安全法规的预训练数据的需求。本文介绍 Common Corpus,这是目前最大的开源语言模型预训练数据集。该数据集中的内容均为无版权或开放授权,总计约两万亿标记符。数据涵盖从高资源欧洲语言到低资源罕见语言的广泛语言种类,并包含大量代码数据。数据来源在领域和时间跨度上的多样性,为跨知识领域的科研与创业应用开辟了路径。本文详细介绍了数据采集来源及过滤与整理过程。我们在 Common Corpus 上训练了两个小型语言模型,结果表明其性能与同规模模型相当,证明该数据集适用于多语言预训练。Common Corpus 对大型语言模型的开放科学生态具有重要贡献。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions of copyrighted or proprietary content, which raises questions about the legal use of such models. This underscores the need for truly open pre-training data that complies with data security regulations. In this paper, we introduce Common Corpus, the largest open dataset for LLM pre-training. The data assembled in Common Corpus are either uncopyrighted or under open licenses, totaling about two trillion tokens. The dataset contains a wide variety of languages, ranging from the high-resource European languages to some low-resource languages rarely represented in pre-training datasets. In addition, it includes a large amount of code data. The diversity of data sources in terms of covered domains and time periods opens up the paths for both research and entrepreneurial needs across diverse areas of knowledge. In this paper, we present the detailed provenance of data assembling and the details of dataset filtering and curation. We train two small language models on Common Corpus and find that they perform comparably to other models of their size, indicating that our dataset is suitable for multilingual pretraining. Common Corpus represents a key contribution to the ecosystem for open science research on Large Language Models.

开源数据语言模型多语言合规训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。