arXiv:2503.13134cs.CV2025-03被引 3

将CLIP与MoCo结合,提升医学影像零样本学习能力

Enhancing zero-shot learning in medical imaging: integrating clip with advanced techniques for improved chest x-ray analysis

  • 融合CLIP与动量对比学习,构建新模型MoCoCLIP
  • 在NIH ChestXray14上比CheXZero提升6.5%准确率
  • 适合缺乏标注数据的医疗AI研究者使用

由于医学影像数据量庞大,需要先进AI方法辅助放射科医生从胸部X光片(CXRs)中诊断胸腔疾病。现有深度学习模型通常依赖大量标注数据,而医学影像因标注耗时且需专家参与,数据稀缺。本文通过将对比语言-图像预训练(CLIP)与动量对比(MoCo)结合,提出MoCoCLIP模型,提升医学影像中的零样本学习能力。该方法有效应对类别不平衡和无标签数据挑战,增强肺部病变检测性能。在NIH ChestXray14数据集上的实验表明,MoCoCLIP相比最先进模型CheXZero实现约6.5%的相对提升;在CheXpert数据集上,平均AUC达0.750,优于CheXZero的0.746,体现出更强的泛化能力。

原文摘要 · Abstract (English)

Due to the large volume of medical imaging data, advanced AI methodologies are needed to assist radiologists in diagnosing thoracic diseases from chest X-rays (CXRs). Existing deep learning models often require large, labeled datasets, which are scarce in medical imaging due to the time-consuming and expert-driven annotation process. In this paper, we extend the existing approach to enhance zero-shot learning in medical imaging by integrating Contrastive Language-Image Pre-training (CLIP) with Momentum Contrast (MoCo), resulting in our proposed model, MoCoCLIP. Our method addresses challenges posed by class-imbalanced and unlabeled datasets, enabling improved detection of pulmonary pathologies. Experimental results on the NIH ChestXray14 dataset demonstrate that MoCoCLIP outperforms the state-of-the-art CheXZero model, achieving relative improvement of approximately 6.5%. Furthermore, on the CheXpert dataset, MoCoCLIP demonstrates superior zero-shot performance, achieving an average AUC of 0.750 compared to CheXZero with 0.746 AUC, highlighting its enhanced generalization capabilities on unseen data.

零样本学习医学影像CLIP胸部X光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。