arXiv:2412.10372cs.CV2024-12被引 64

构建首个跨六类医学影像的通用图文预训练模型,解决数据稀缺难题。

UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities

  • 用大语言模型自动构建530万组医学影像-文本对,覆盖六类模态。
  • 在21个数据集上零样本性能超越现有模型,用3倍少数据实现更高精度。
  • 开源数据集与代码,推动跨模态医学视觉研究发展。

基于对比学习的视觉语言模型在自然图像任务中取得显著进展,但在医学领域应用受限,主要因缺乏公开、大规模的医学图像-文本数据集。现有医学VLM多基于封闭源或较小的开源数据集,泛化能力差,且通常仅针对单一或有限模态,难以跨模态应用。为此,我们构建了UniMed,一个大规模、开源的多模态医学数据集,包含超过530万组图像-文本对,覆盖六类成像模态:X光、CT、MRI、超声、病理和眼底图像。UniMed通过利用大语言模型将各模态特定分类数据转换为图文格式,并融合现有医学图文数据,实现了可扩展的VLM预训练。基于此,我们训练了UniMed-CLIP,一种面向六类模态的统一视觉语言模型,在21个数据集上的零样本评估中显著优于现有通用模型,且性能媲美专用医学模型。例如,相比在专有数据上训练的BiomedCLIP,UniMed-CLIP平均提升+12.61,同时仅使用其1/3的训练数据。为促进后续研究,我们已在GitHub开源UniMed数据集、训练代码与模型。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) trained via contrastive learning have achieved notable success in natural image tasks. However, their application in the medical domain remains limited due to the scarcity of openly accessible, large-scale medical image-text datasets. Existing medical VLMs either train on closed-source proprietary or relatively small open-source datasets that do not generalize well. Similarly, most models remain specific to a single or limited number of medical imaging domains, again restricting their applicability to other modalities. To address this gap, we introduce UniMed, a large-scale, open-source multi-modal medical dataset comprising over 5.3 million image-text pairs across six diverse imaging modalities: X-ray, CT, MRI, Ultrasound, Pathology, and Fundus. UniMed is developed using a data-collection framework that leverages Large Language Models (LLMs) to transform modality-specific classification datasets into image-text formats while incorporating existing image-text data from the medical domain, facilitating scalable VLM pretraining. Using UniMed, we trained UniMed-CLIP, a unified VLM for six modalities that significantly outperforms existing generalist VLMs and matches modality-specific medical VLMs, achieving notable gains in zero-shot evaluations. For instance, UniMed-CLIP improves over BiomedCLIP (trained on proprietary data) by an absolute gain of +12.61, averaged over 21 datasets, while using 3x less training data. To facilitate future research, we release UniMed dataset, training codes, and models at https://github.com/mbzuai-oryx/UniMed-CLIP.

医学影像图文预训练多模态开源数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。