开源100万亿tokens数据集,助力透明高效语言模型训练
RedPajama: an Open Dataset for Training Large Language Models
- 复现LLaMA数据集并发布纯网络原始数据集RedPajama-V2
- 包含超100万亿标记的多领域数据,附质量信号与元数据
- 支持模型透明训练,已用于Snowflake Arctic等生产级模型
大型语言模型已成为人工智能、科学及社会的核心技术,但其数据集构建与筛选策略仍不明确。多数高性能模型缺乏数据采集与模型开发过程的透明度,阻碍了完全开源语言模型的发展。本文指出三大关键数据挑战:模型开发透明性、高质量数据获取、数据清洗与分析所需的工具与元数据可用性。为此,我们发布RedPajama-V1,即LLaMA训练数据集的开源复现版本;同时推出RedPajama-V2,一个由原始未过滤网络文本构成的超大规模数据集,附带质量信号和元数据。两个数据集合计超过100万亿个标记,覆盖多个领域,其质量信号有助于数据筛选,旨在推动新数据集的开发。目前,这些数据集已被用于Snowflake Arctic、Salesforce XGen和AI2 OLMo等生产级语言模型的训练。通过针对解码器仅有的语言模型(最大1.6B参数)进行分析与消融实验,我们验证了网页数据质量信号的有效性,证明其可有效筛选出高质量子集,凸显RedPajama在推动可解释、高性能语言模型规模化发展方面的潜力。
原文摘要 · Abstract (English)
Large language models are increasingly becoming a cornerstone technology in artificial intelligence, the sciences, and society as a whole, yet the optimal strategies for dataset composition and filtering remain largely elusive. Many of the top-performing models lack transparency in their dataset curation and model development processes, posing an obstacle to the development of fully open language models. In this paper, we identify three core data-related challenges that must be addressed to advance open-source language models. These include (1) transparency in model development, including the data curation process, (2) access to large quantities of high-quality data, and (3) availability of artifacts and metadata for dataset curation and analysis. To address these challenges, we release RedPajama-V1, an open reproduction of the LLaMA training dataset. In addition, we release RedPajama-V2, a massive web-only dataset consisting of raw, unfiltered text data together with quality signals and metadata. Together, the RedPajama datasets comprise over 100 trillion tokens spanning multiple domains and with their quality signals facilitate the filtering of data, aiming to inspire the development of numerous new datasets. To date, these datasets have already been used in the training of strong language models used in production, such as Snowflake Arctic, Salesforce's XGen and AI2's OLMo. To provide insight into the quality of RedPajama, we present a series of analyses and ablation studies with decoder-only language models with up to 1.6B parameters. Our findings demonstrate how quality signals for web data can be effectively leveraged to curate high-quality subsets of the dataset, underscoring the potential of RedPajama to advance the development of transparent and high-performing language models at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。