打造13200万份无版权风险的LLM训练数据,推动AI合法发展。
The KL3M Data Project: Copyright-Clean Training Resources for Large Language Models
- 构建16个来源的13200万文档数据集,严格通过版权合规验证。
- 包含万亿级标记、原始文档、标注数据和分词表示等完整资源链。
- 开源所有工具与数据,适合关注法律合规与可持续AI的研究者使用。
几乎所有大型语言模型的预训练数据都存在全球范围内的版权侵权和合同违约潜在风险,给用户和开发者带来法律不确定性。KL3M数据项目直接应对这一关键问题,推出迄今为止规模最大、覆盖最全的训练数据流水线,最大限度降低版权或合同违约风险。该项目基于超过1.32亿份文档、跨越万亿级标记的数据,涵盖16个经严格验证符合本研究详细版权与许可协议的来源。我们公开发布整个流水线,包括:1)获取与处理文档的源代码;2)原始文档格式及出处、元数据;3)标准化提取内容;4)预分词表示;5)多种中后训练资源,如问答、摘要、改写、写作、分类、预测和对话数据。所有资源均以CC-BY协议在S3、Hugging Face和GitHub免费开放。我们致力于持续推进更伦理、合法且可持续的AI模型研发与应用。
原文摘要 · Abstract (English)
Practically all large language models have been pre-trained on data that is subject to global uncertainty related to copyright infringement and breach of contract. This creates potential risk for users and developers due to this uncertain legal status. The KL3M Data Project directly confronts this critical issue by introducing the largest comprehensive training data pipeline that minimizes risks related to copyright or breach of contract. The foundation of this project is a corpus of over 132 million documents and trillions of tokens spanning 16 different sources that have been verified to meet the strict copyright and licensing protocol detailed herein. We are releasing the entire pipeline, including 1) the source code to acquire and process these documents, 2) the original document formats with associated provenance and metadata, 3) extracted content in a standardized format, 4) pre-tokenized representations of the documents, and 5) various mid- and post-train resources such as question-answer, summarization, conversion, drafting, classification, prediction, and conversational data. All of these resources are freely available to the public on S3, Hugging Face, and GitHub under CC-BY terms. We are committed to continuing this project in furtherance of a more ethical, legal, and sustainable approach to the development and use of AI models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。