arXiv:2411.07715cs.AI2024-11被引 2

梳理大模型训练数据现状,助你快速掌握核心资源与方法

Training Data for Large Language Model

  • 系统总结预训练与微调数据的规模、采集方式和处理流程
  • 覆盖主流数据类型特征,提供可直接使用的开源数据集清单
  • 适合想构建高质量数据集的研究者与开发者参考

2022年,随着ChatGPT的发布,大规模语言模型引起广泛关注。ChatGPT不仅在参数量和预训练语料规模上超越以往模型,更通过在海量高质量人工标注数据上微调,实现了革命性性能提升。这一进展使企业和研究机构认识到,打造更智能、更强的大模型依赖于丰富且高质量的数据集。因此,数据集的构建与优化已成为人工智能领域的关键焦点。本文总结了当前大模型预训练与微调数据的现状,涵盖数据规模、采集方法、数据类型与特征、处理流程,并对现有开源数据集进行了综述。

原文摘要 · Abstract (English)

In 2022, with the release of ChatGPT, large-scale language models gained widespread attention. ChatGPT not only surpassed previous models in terms of parameters and the scale of its pretraining corpus but also achieved revolutionary performance improvements through fine-tuning on a vast amount of high-quality, human-annotated data. This progress has led enterprises and research institutions to recognize that building smarter and more powerful models relies on rich and high-quality datasets. Consequently, the construction and optimization of datasets have become a critical focus in the field of artificial intelligence. This paper summarizes the current state of pretraining and fine-tuning data for training large-scale language models, covering aspects such as data scale, collection methods, data types and characteristics, processing workflows, and provides an overview of available open-source datasets.

大模型训练数据集预训练开源数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。