arXiv:2503.19801cs.CVcs.AI2025-03被引 3

用报告相似性增强医学影像预训练,提升小样本下的多模态分析能力。

SeLIP: Similarity Enhanced Contrastive Language Image Pretraining for Multi-modal Head MRI

  • 引入文本语法与语义相似性匹配,减少对海量数据的依赖。
  • 在图像-文本检索、分类、分割任务中表现优异,验证模型有效性。
  • 适合缺乏标注数据的医疗AI研究者,尤其关注多模态建模场景。

尽管深度学习在众多医学图像分析任务中展现出巨大潜力,但受限于人工标注数据不足,其实际应用仍受制约。鉴于临床放射学检查通常附带描述图像的放射学报告,本文提出一种基于图像与对应放射学发现的对比学习框架,构建头部MRI多模态基础模型。具体地,设计了一种融合语法与语义相似性的混合匹配度量,以缓解传统对比学习对超大规模数据集的依赖。所提出的相似性增强对比语言图像预训练(SeLIP)能够有效提取更具判别力的特征。实验表明,该方法在图像-文本检索、分类及图像分割等下游任务中均表现良好,凸显了在构建医学图像基础模型时考虑不同图像描述间相似性的重要性。

原文摘要 · Abstract (English)

Despite that deep learning (DL) methods have presented tremendous potential in many medical image analysis tasks, the practical applications of medical DL models are limited due to the lack of enough data samples with manual annotations. By noting that the clinical radiology examinations are associated with radiology reports that describe the images, we propose to develop a foundation model for multi-model head MRI by using contrastive learning on the images and the corresponding radiology findings. In particular, a contrastive learning framework is proposed, where a mixed syntax and semantic similarity matching metric is integrated to reduce the thirst of extreme large dataset in conventional contrastive learning framework. Our proposed similarity enhanced contrastive language image pretraining (SeLIP) is able to effectively extract more useful features. Experiments revealed that our proposed SeLIP performs well in many downstream tasks including image-text retrieval task, classification task, and image segmentation, which highlights the importance of considering the similarities among texts describing different images in developing medical image foundation models.

多模态医学影像对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。