构建了250亿词元的医学语料库,填补互联网医疗文本数据空白。
TheBlueScrubs-v1, a comprehensive curated medical dataset derived from the internet
- 从互联网海量文本中筛选医学内容,两阶段过滤确保质量
- 含110亿肿瘤相关词元,支持精准医学研究
- 适合医疗大模型训练与安全评估,临床验证效果佳
为解决当前公开医学数据集规模和范围不足的问题,我们构建了TheBlueScrubs-v1,一个涵盖超过250亿医疗词元的高质量语料库,规模接近PubMed的三倍。该数据源自大规模互联网语料,通过两阶段过滤流程:首先用逻辑回归模型(外部验证AUC≈0.95)筛选文档,再由700亿参数的Llama 3.1指令模型进行人工级验证。每条文本获得三项基于大模型的评分:医学相关性、精确度与事实细节、安全性与伦理标准。临床医生评审显示与自动化评估高度一致;另有专用癌症分类器标注约110亿肿瘤相关词元。两个应用演示表明:将安全评估结果蒸馏至小型BERT类模型,其在未见数据上达到AUC近0.96;在过滤子集上微调轻量级大模型,在公开及私有医学基准测试中均显著优于基线。本数据描述详细阐述了数据构建与验证过程,凸显其在医疗AI研究中的潜力。
原文摘要 · Abstract (English)
The need for robust and diverse data sets to train clinical large language models (cLLMs) is critical given that currently available public repositories often prove too limited in size or scope for comprehensive medical use. While resources like PubMed provide foundational medical literature, they capture only a narrow range of formal publications and omit the broader medical discourse on the internet. To address these deficits, we introduce TheBlueScrubs-v1, a curated dataset of over 25 billion medical tokens - nearly three times larger than PubMed - drawn from a broad-scale internet corpus. Our two-stage filtering pipeline employs a Logistic Regression model for document screening (achieving an AUC of approximately 0.95 on external validation), followed by verification via a 70B-parameter Llama 3.1 instruct model. Each text is assigned three LLM-based quality scores encompassing medical relevance, precision and factual detail, and safety and ethical standards. Clinician reviews confirm high concordance with these automated evaluations, and a specialized cancer classifier further labels approximately 11 billion oncology tokens. Two demonstration tasks highlight the dataset's practical value: first, we distill the safety evaluations to a smaller BERT-style model that reaches an AUC near 0.96 on unseen data; second, we fine-tune a compact LLM on a filtered subset, showing measurable improvements over standard baselines in medical benchmarks as well as private ones. This Data Descriptor details the dataset's creation and validation, underscoring its potential utility for medical AI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。