用大模型标注40万篇医学文献,构建高质量临床文本数据集。
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
- 用大模型先标注40万段文献,再用小模型扩散标签到全量PMC-OA数据。
- 提取出200万临床病例段落,其中45万条为高教育质量内容。
- 适合医学NLP研究者,尤其关注临床文本和高效预训练方法的人。
我们提出Biomed-Enriched,一个通过两阶段标注流程从PubMed构建的生物医学文本数据集。第一阶段,大语言模型对40万篇科学文章中的段落进行标注,评估其类型(综述、研究、临床案例等)、领域(临床、生物医学等)及教育质量(1-5分),衡量其对大学水平学习的实用性。这些标注用于微调小型语言模型,进而将标签传播至整个PMC-OA语料库。由此可提取出包含200万临床病例段落的数据子集,其中超过45万条来自可商用文章,且质量较高;并可通过质量过滤与领域上采样构建多种变体。由于医院记录受隐私限制难以公开,该数据集提供了大规模、开放获取的临床病例资源,是生物医学与临床NLP的重要工具。初步持续预训练实验表明,针对临床内容的上采样使OLMo2在MMLU ProfMed上性能提升约5%,教育质量过滤使MedQA和MedMCQA提升约1%。结合两种策略可实现更快收敛,仅需三分之一训练样本即达相同性能,显示出更高效、精准的生物医学预训练潜力。
原文摘要 · Abstract (English)
We introduce Biomed-Enriched, a biomedical text dataset constructed from PubMed via a two-stage annotation process. In the first stage, a large language model annotates 400K paragraphs from PubMed scientific articles, assigning scores for their type (review, study, clinical case, other), domain (clinical, biomedical, other), and educational quality. The educational quality score (rated 1 to 5) estimates how useful a paragraph is for college-level learning. These annotations are then used to fine-tune a small language model, which propagates the labels across the full PMC-OA corpus. The resulting metadata allows us to extract refined subsets, including 2M clinical case paragraphs with over 450K high-quality ones from articles with commercial-use licenses, and to construct several variants via quality filtering and domain upsampling. Clinical text is typically difficult to access due to privacy constraints, as hospital records cannot be publicly shared. Hence, our dataset provides an alternative large-scale, openly available collection of clinical cases from PubMed, making it a valuable resource for biomedical and clinical NLP. Preliminary continual-pretraining experiments with OLMo2 suggest these curated subsets enable targeted improvements, with clinical upsampling boosting performance by ~5% on MMLU ProfMed and educational quality filtering improving MedQA and MedMCQA by ~1%. Combinations of these techniques led to faster convergence, reaching same performance with a third of training tokens, indicating potential for more efficient and effective biomedical pretraining strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。