用自监督对比学习让ViT更好识别心脏MRI序列
Self-Supervised Contrastive Learning for Cardiac MR Sequence Classification

- 用图像级自监督对比学习微调ViT,提升医学影像适配性
- 在四种常见心脏MRI序列上分类AUC超0.75,优于传统方法
- 适合医疗影像领域研究者,尤其关注模型迁移与小样本学习
视觉变换器(ViT)模型凭借自注意力机制在各类视觉任务中表现出强大的泛化能力,包括图像分类。然而,这些模型通常在通用公共数据集上预训练,缺乏医学成像应用所需的领域知识。本研究探讨了将ViT模型应用于心脏磁共振(MR)图像的适应性,使用内部构建的数据集进行实验。我们发现,预训练的ViT特征无法有效迁移到心脏MR领域。为克服此限制,我们提出一种基于图像的自监督对比学习适应策略,其性能显著优于传统监督训练方法。此外,经调整的ViT模型在外部MR数据集(如BraTS和ADNI)上也展现出强泛化能力。通过消融实验,我们进一步研究了批量大小和数据集规模对性能的影响。最终,该适配模型在四种最常见的心脏MR序列上的分类AUC均超过0.75。
原文摘要 · Abstract (English)
Vision Transformer (ViT) models, utilizing self-attention mechanisms, have demonstrated robust generalization capabilities across various vision tasks, including image classification. However, these models, typically pretrained on general public datasets, often lack the specialized domain knowledge necessary for medical imaging applications. In this study, we investigate the adaptation of ViT models, specifically for cardiac magnetic resonance (MR) images, using an in-house dataset. We found that pretrained ViT features do not effectively transfer to the cardiac MR domain. To overcome this limitation, we introduce an adaptation strategy that utilizes image-based self-supervised contrastive learning, demonstrating superior performance compared to traditional supervised training approaches. Moreover, our adapted ViT model exhibits strong generalization to external MR datasets such as BraTS and ADNI. Through ablation studies, we further investigate the impact of batch size and dataset scale on performance. Ultimately, our adapted model achieves classification AUC exceeding 0.75 across the four most common cardiac MR sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。