arXiv:2507.15275cs.CL2025-07

构建覆盖中西医的204万汉字医学数据集,支持大模型预训练与强化学习。

ChiMed 2.0: Advancing Chinese Medical Dataset in Facilitating Large Language Modeling

  • 整合线上平台与LLM生成数据,覆盖中医典籍与现代医学内容。
  • 含164.8K预训练文档、351.6K问答对和41.7K偏好数据,支持全链路训练。
  • 在多个模型规模上验证有效,适合中文医疗大模型研发者使用。

构建高质量数据资源对推动特定领域人工智能研究与应用至关重要,尤其在中文医学领域。现有中文医学数据集规模有限且领域覆盖狭窄,难以满足有效预训练所需的多样化语料。此外,多数数据集仅用于大模型微调,不支持预训练与人类反馈强化学习(RLHF)。本文提出中文医学数据集ChiMed 2.0,扩展自先前工作ChiMed,涵盖从中文医学网络平台收集及由大语言模型生成的数据。该数据集共包含204.4万汉字,覆盖传统中医经典与现代通用医学内容,其中包含164.8K用于预训练的文档、351.6K用于监督微调(SFT)的问答对,以及41.7K用于强化学习(RLHF)的偏好数据对。为验证该数据集在训练中文医学大模型方面的有效性,我们在代表性通用领域大模型上开展预训练、SFT与RLHF实验,并在医学基准数据集上评估性能。结果表明,在不同模型规模下均取得性能提升,验证了该数据集的有效性与适用性。

原文摘要 · Abstract (English)

Building high-quality data resources is crucial for advancing artificial intelligence research and applications in specific domains, particularly in the Chinese medical domain. Existing Chinese medical datasets are limited in size and narrow in domain coverage, falling short of the diverse corpora required for effective pre-training. Moreover, most datasets are designed solely for LLM fine-tuning and do not support pre-training and reinforcement learning from human feedback (RLHF). In this paper, we propose a Chinese medical dataset named ChiMed 2.0, which extends our previous work ChiMed, and covers data collected from Chinese medical online platforms and generated by LLMs. ChiMed 2.0 contains 204.4M Chinese characters covering both traditional Chinese medicine classics and modern general medical data, where there are 164.8K documents for pre-training, 351.6K question-answering pairs for supervised fine-tuning (SFT), and 41.7K preference data tuples for RLHF. To validate the effectiveness of our approach for training a Chinese medical LLM, we conduct further pre-training, SFT, and RLHF experiments on representative general domain LLMs and evaluate their performance on medical benchmark datasets. The results show performance gains across different model scales, validating the dataset's effectiveness and applicability.

医学大模型中文数据集LLM训练预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。