arXiv:2607.07673cs.CVcs.LG2026-07

从610万篇医学文献中自动构建1100万对高质量图文数据,提升医疗多模态模型性能。

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

论文配图:MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
图 1 · 摘自论文原文
  • 自动化框架将开源文献转化为高保真医学图文对,持续更新
  • 在26个基准上使零样本分类准确率提升7.1个百分点,仅用一半数据量
  • 适合医疗视觉问答、皮肤病图像检索等临床场景应用

医学本质为多模态,需整合多种数据流信息。然而多模态基础模型发展受限于大规模高质量临床数据的匮乏。尽管文献库PubMed Central(PMC)提供专家撰写图像-文本数据,现有基于PMC的数据集仍存在保真度低、可复现性差、缺乏临床验证等问题。我们提出MedPMC,一个自动化、可持续更新的框架,将授权宽松的文献转化为医疗多模态模型的高保真基础设施。应用于610万篇PMC文章,生成1100万对医学图像-文本对。组件评估显示:初步筛选F1=93.2,多面板图检测F1=96.5,图分离mAP=89.8,标题分离与对齐F1=81.4;ROUGE-L=85.3,医学图像分类F1=96.5。五名标注员(三人具医学背景)人工审查发现,MedPMC图像中95.3%具有医学相关性,远高于此前数据集的19.7%。在覆盖11个专科的26项基准测试中,经MedPMC训练的CLIP类模型,零样本平均AUC比最强基线提升7.1个百分点,且所用图像-文本对不足其一半。作为多模态大语言模型的视觉编码器,在两项医疗视觉问答任务中分别提升1.9和16.9个百分点。在10,524张耶鲁纽黑文医疗系统皮肤科照片中,形态到图像检索的Recall@5提升11.7个百分点。结果表明,高保真文献整理能显著增强医疗多模态基础模型在基准与真实临床场景中的表现。我们公开发布该框架、语料库、基准与预训练模型。

原文摘要 · Abstract (English)

Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundation models is constrained by limited access to large-scale, high-quality clinical data. Although PubMed Central (PMC) offers a complementary source of expert-authored image-text data, existing PMC-derived resources remain limited in fidelity, reproducibility, and clinical validation. We introduce MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models. Applied to 6.1 million PMC articles, MedPMC curated 11 million medical image-text pairs. Component evaluations showed strong performance for initial screening (F1 = 93.2), multi-panel figure detection (F1 = 96.5), figure separation (mAP = 89.8), caption separation and alignment (F1 = 81.4; ROUGE-L = 85.3), and medical figure classification (F1 = 96.5). Manual review by five annotators, three with medical training, found 95.3% of MedPMC images medically relevant, versus 19.7% in a prior PMC-derived dataset. Across 26 benchmarks spanning 11 specialties, a MedPMC-trained CLIP-style model improved average zero-shot AUC by 7.1 percentage points over the strongest architecture-matched biomedical CLIP baseline despite using fewer than half as many image-text pairs. As the vision encoder in a multimodal large language model, it improved medical visual question-answering by 1.9 and 16.9 percentage points across two benchmarks. In 10,524 Yale New Haven Health System dermatology photographs, it improved morphology-to-image retrieval Recall@5 by 11.7 percentage points. These findings show that high-fidelity literature curation strengthens medical multimodal foundation models across benchmark and clinical settings. We publicly release the framework, corpus, benchmarks, and pretrained models.

医疗多模态图文对数据构建基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。