统一自监督框架提升医学图文预训练效果
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training
- 融合跨模态、单模态与融合模态的自监督学习
- 在图像检索、分类和视觉问答任务上超越现有方法
- 特别适配医学影像特性,缓解数据稀缺问题
基于对比学习的视觉语言预训练近期显著提升了计算机视觉任务性能。然而,在医学领域,由于隐私、敏感性和标注复杂性,获取多模态数据成本高昂且困难。为缓解数据稀缺并提升模型表现,我们提出针对医学视觉语言预训练的统一自监督框架 Uni-Mlip。该框架在数据层与特征层无缝整合跨模态、单模态及融合模态自监督技术,并针对医学图像特点定制单模态图像自监督策略。在不同规模数据集上的实验表明,Uni-Mlip 在图像-文本检索、图像分类和视觉问答三项关键下游任务中显著优于当前最优方法。
原文摘要 · Abstract (English)
Recent advancements in vision-language pre-training via contrastive learning have significantly improved performance across computer vision tasks. However, in the medical domain, obtaining multimodal data is often costly and challenging due to privacy, sensitivity, and annotation complexity. To mitigate data scarcity while boosting model performance, we introduce \textbf{Uni-Mlip}, a unified self-supervision framework specifically designed to enhance medical vision-language pre-training. Uni-Mlip seamlessly integrates cross-modality, uni-modality, and fused-modality self-supervision techniques at the data-level and the feature-level. Additionally, Uni-Mlip tailors uni-modal image self-supervision to accommodate the unique characteristics of medical images. Our experiments across datasets of varying scales demonstrate that Uni-Mlip significantly surpasses current state-of-the-art methods in three key downstream tasks: image-text retrieval, image classification, and visual question answering (VQA).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。