arXiv:2503.14377eess.IVcs.CV2025-03被引 7

高质量医学图文数据能显著提升多模态模型性能。

Advancing Medical Representation Learning Through High-Quality Data

  • 构建220万对医学图文数据,含图像模态标注和文本引用信息。
  • 在检索与零样本分类任务中,小而高质量数据优于大而低质数据。
  • 适合医疗影像与自然语言处理交叉研究者使用。

尽管医学视觉-语言数据集规模不断增长,但数据质量对模型性能的影响仍缺乏深入研究。本文提出Open-PMC,一个来自PubMed Central的高质量医学数据集,包含220万张图像-文本对,附带图像模态标注、子图信息及摘要性文本引用。特别地,文本引用提供了比传统标题更丰富的医学上下文。通过大量实验,我们在检索与零样本分类任务中对比了Open-PMC与更大规模数据集的表现。结果表明,数据质量——而非单纯规模——是性能提升的关键。我们进一步分析了特征表示,强调数据精炼质量在推动多模态医学AI发展中的核心作用。本文发布Open-PMC数据集、训练模型及代码库。

原文摘要 · Abstract (English)

Despite the growing scale of medical Vision-Language datasets, the impact of dataset quality on model performance remains under-explored. We introduce Open-PMC, a high-quality medical dataset from PubMed Central, containing 2.2 million image-text pairs, enriched with image modality annotations, subfigures, and summarized in-text references. Notably, the in-text references provide richer medical context, extending beyond the abstract information typically found in captions. Through extensive experiments, we benchmark Open-PMC against larger datasets across retrieval and zero-shot classification tasks. Our results show that dataset quality-not just size-drives significant performance gains. We complement our benchmark with an in-depth analysis of feature representation. Our findings highlight the crucial role of data curation quality in advancing multimodal medical AI. We release Open-PMC, along with the trained models and our codebase.

医学多模态数据质量视觉语言Open-PMC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。