构建1800万对高质量医学图文数据集,提升医疗视觉语言模型性能
Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning
- 用Transformer提取图表子图与子标题,结合文献引用扩充上下文
- 建成1800万对图文数据,覆盖放射、显微和可见光医学图像
- 在6项检索和19项零样本分类任务中刷新医疗表征学习新纪录
在生物医学视觉-语言建模中,数据集通常从科学文献中挖掘,图像与文字配对存在简短、依赖上下文且信息不全的问题。现有子图提取工作在规模和泛化能力上均有限,且缺乏丰富的医学上下文。本文重新审视数据构建这一基础环节,提出融合Transformer子图检测、子标题提取及基于内联参考的上下文文本增强的流程。所提子图提取模型在50万张复合图上训练,在真实与合成基准上达到领先水平。基于该流程,我们构建并发布Open-PMC-18M,一个包含1800万对高保真医学图文对的大规模数据集,涵盖放射、显微和可见光摄影。我们在该数据集上训练视觉-语言模型,并在6项检索与19项零样本分类任务上进行评估,涉及三种主要模态。结果表明,基于本数据集训练的模型在医学表征学习中达到新标杆。我们公开数据集、模型与代码,以支持可复现基准与进一步研究。
原文摘要 · Abstract (English)
In biomedical vision-language modeling, datasets are typically mined from scientific literature, pairing compound figures with captions that are short, context-dependent, and oftern partially informative. Prior work on subfigure extraction has been limited in both dataset size and generalizability. In addition, no existing effort has incorporated rich medical context in image-text pairs. We revisit data curation as a foundational component of effective biomedical representation learning. Our data curation process integrates transformer-based subfigure detection, subcaption extraction, and contextual text enrichment derived from inline references. Our subfigure extraction model, trained on a corpus of 500,000 compound figures, achieves state-of-the-art performance on real and synthetic benchmarks. Using this process, we curate and release Open-PMC-18M, a large-scale high-fidelity biomedical dataset comprising 18 million image-text pairs, spanning radiology, microscopy, and visible light photography. We train vision-language models on our dataset and perform extensive evaluation on 6 retrieval and 19 zero-shot classification tasks across three major modalities. The models trained on our dataset set a new state-of-the-art results in medical representation learning. We release our dataset, models, and code to support reproducible benchmarks and further study into biomedical vision-language modeling and representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。