用3600万张心脏核磁图训练通用模型,提升各类任务表现
Towards a vision foundation model for comprehensive assessment of Cardiac MRI
- 自监督预训练+微调,统一处理9类心脏核磁任务
- 在小样本下仍优于现有方法,多数任务达顶尖水平
- 适合医疗影像研究者,尤其数据少的场景
心脏磁共振(CMR)是评估心脏形态与功能的金标准,但其图像分析任务多样且复杂,深度学习模型常因标注数据稀缺而难以训练。本文提出一种针对CMR的视觉基础模型,在3600万张未标注的CMR图像上进行自监督预训练,随后在9个典型临床任务中进行监督微调,涵盖分类、分割、关键点定位和病灶检测。实验表明,该模型在不同标注数据量下均表现出更高的准确率与鲁棒性,尤其在少样本学习场景下优势显著。其零样本性能已接近当前最优水平,为资源有限的医学影像分析提供了高效统一的解决方案。
原文摘要 · Abstract (English)
Cardiac magnetic resonance imaging (CMR), considered the gold standard for noninvasive cardiac assessment, is a diverse and complex modality requiring a wide variety of image processing tasks for comprehensive assessment of cardiac morphology and function. Advances in deep learning have enabled the development of state-of-the-art (SoTA) models for these tasks. However, model training is challenging due to data and label scarcity, especially in the less common imaging sequences. Moreover, each model is often trained for a specific task, with no connection between related tasks. In this work, we introduce a vision foundation model trained for CMR assessment, that is trained in a self-supervised fashion on 36 million CMR images. We then finetune the model in supervised way for 9 clinical tasks typical to a CMR workflow, across classification, segmentation, landmark localization, and pathology detection. We demonstrate improved accuracy and robustness across all tasks, over a range of available labeled dataset sizes. We also demonstrate improved few-shot learning with fewer labeled samples, a common challenge in medical image analyses. We achieve an out-of-box performance comparable to SoTA for most clinical tasks. The proposed method thus presents a resource-efficient, unified framework for CMR assessment, with the potential to accelerate the development of deep learning-based solutions for image analysis tasks, even with few annotated data available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。