arXiv:2412.03398cs.CL2024-12被引 6

用网页数据构建通用与专业领域的高质量训练集

RedStone: Curating General, Code, Math, and QA Data for Large Language Models

  • 通过自动化管道从Common Crawl提取数据,无需人工标注
  • 生成覆盖语言、代码、数学和问答的多领域训练数据集
  • 适合想低成本构建垂直领域模型的研究者使用

在高质量、精心筛选的数据集上预训练大语言模型,被广泛认为是提升其性能和泛化能力的关键。本研究探索了公共网络抓取数据(Common Crawl)作为全面且灵活资源的潜力,用于大语言模型的预训练,涵盖通用语言理解与特定领域知识。我们提出RedStone,一种创新且可扩展的数据处理管道,能够从Common Crawl中提取并处理数据,支持构建大规模、多样化的预训练数据集。相比传统依赖昂贵人工标注和领域专业知识的数据集,RedStone利用Common Crawl的广度,实现多领域定制化数据生产。本文展示了其在通用语言、代码、数学和问答任务等领域的应用能力。其灵活性使其易于拓展至其他专业领域,显著降低创建有价值领域数据集的门槛。结果表明,通过有效管道如RedStone,Common Crawl可成为丰富且可再生的预训练数据来源,为大语言模型的领域适配与知识发现开辟新路径。本工作强调了创新数据获取策略的重要性,并凸显了网络规模数据在大语言模型持续演进中的关键作用。RedStone代码与数据样本将公开发布于\url{https://aka.ms/redstone}。

原文摘要 · Abstract (English)

Pre-training Large Language Models (LLMs) on high-quality, meticulously curated datasets is widely recognized as critical for enhancing their performance and generalization capabilities. This study explores the untapped potential of Common Crawl as a comprehensive and flexible resource for pre-training LLMs, addressing both general-purpose language understanding and specialized domain knowledge. We introduce RedStone, an innovative and scalable pipeline engineered to extract and process data from Common Crawl, facilitating the creation of extensive and varied pre-training datasets. Unlike traditional datasets, which often require expensive curation and domain-specific expertise, RedStone leverages the breadth of Common Crawl to deliver datasets tailored to a wide array of domains. In this work, we exemplify its capability by constructing pre-training datasets across multiple fields, including general language understanding, code, mathematics, and question-answering tasks. The flexibility of RedStone allows for easy adaptation to other specialized domains, significantly lowering the barrier to creating valuable domain-specific datasets. Our findings demonstrate that Common Crawl, when harnessed through effective pipelines like RedStone, can serve as a rich, renewable source of pre-training data, unlocking new avenues for domain adaptation and knowledge discovery in LLMs. This work also underscores the importance of innovative data acquisition strategies and highlights the role of web-scale data as a powerful resource in the continued evolution of LLMs. RedStone code and data samples will be publicly available at \url{https://aka.ms/redstone}.

数据构建大模型代码生成通用模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。