开源高质量中文语料库,助力大模型训练性能提升
OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training
- 构建多类型中文数据集,覆盖网页、教材、对话等场景
- 在C-Eval任务上显著提升小参数模型表现
- 适合中文大模型预训练与微调,支持可复现数据流程
大型语言模型(LLMs)表现出色,但其性能高度依赖预训练语料质量。针对中文领域高质量数据稀缺的问题,我们提出OpenCSG中文语料库,包含Fineweb-edu-chinese、Fineweb-edu-chinese-v2、Cosmopedia-chinese和Smoltalk-chinese四类数据集。其中,Fineweb-edu系列聚焦从多样化中文网络来源过滤出的高质量内容;Cosmopedia-chinese提供合成的教材风格知识型数据;Smoltalk-chinese强调多样化的对话格式数据。该语料库具备高质量文本、跨领域广泛覆盖及可扩展、可复现的数据处理流程。我们通过大量实验验证,包括在小参数模型上的评估,在C-Eval等任务中展现出显著性能提升,证明了其在中文LLM训练中的有效性。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable capabilities, but their success heavily relies on the quality of pretraining corpora. For Chinese LLMs, the scarcity of high-quality Chinese datasets presents a significant challenge, often limiting their performance. To address this issue, we propose the OpenCSG Chinese Corpus, a series of high-quality datasets specifically designed for LLM pretraining, post-training, and fine-tuning. This corpus includes Fineweb-edu-chinese, Fineweb-edu-chinese-v2, Cosmopedia-chinese, and Smoltalk-chinese, each with distinct characteristics: Fineweb-edu datasets focus on filtered, high-quality content derived from diverse Chinese web sources; Cosmopedia-chinese provides synthetic, textbook-style data for knowledge-intensive training; and Smoltalk-chinese emphasizes stylistic and diverse chat-format data. The OpenCSG Chinese Corpus is characterized by its high-quality text, diverse coverage across domains, and scalable, reproducible data curation processes. Additionally, we conducted extensive experimental analyses, including evaluations on smaller parameter models, which demonstrated significant performance improvements in tasks such as C-Eval, showcasing the effectiveness of the corpus for training Chinese LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。