用多模态MRI预训练提升医学图像分析,显著改善分割与分类性能。
Multi-modal Vision Pre-training for Medical Image Analysis
- 设计三种代理任务学习跨模态图像关联,利用240万张脑部MRI数据
- 在6个分割任务上Dice分数提升0.28%至14.47%,4个分类任务准确率增0.65%~18.07%
- 适合需要多模态医学影像建模的研究者和临床辅助诊断系统开发者
自监督学习通过降低真实数据标注需求,显著推动了医学图像分析的发展。当前主流方法主要依赖单模态图像的自监督学习,忽略了跨模态间的重要关联,尤其在自然成组的多模态数据(如同一患者不同功能成像协议的多参数MRI)中尤为明显。为弥补这一缺陷,本文提出一种新型多模态图像预训练方法,包含三个代理任务:跨模态图像重建、模态感知对比学习和模态模板蒸馏,基于超过240万张图像(来自3,755名患者的16,022次扫描)的多模态脑部MRI数据进行训练。为验证模型泛化能力,我们在十个下游任务上进行了广泛实验。结果表明,该方法在六个分割基准上相比现有最优预训练方法,Dice Score提升0.28%–14.47%,在四个独立图像分类任务中准确率持续提升0.65%–18.07%。
原文摘要 · Abstract (English)
Self-supervised learning has greatly facilitated medical image analysis by suppressing the training data requirement for real-world applications. Current paradigms predominantly rely on self-supervision within uni-modal image data, thereby neglecting the inter-modal correlations essential for effective learning of cross-modal image representations. This limitation is particularly significant for naturally grouped multi-modal data, e.g., multi-parametric MRI scans for a patient undergoing various functional imaging protocols in the same study. To bridge this gap, we conduct a novel multi-modal image pre-training with three proxy tasks to facilitate the learning of cross-modality representations and correlations using multi-modal brain MRI scans (over 2.4 million images in 16,022 scans of 3,755 patients), i.e., cross-modal image reconstruction, modality-aware contrastive learning, and modality template distillation. To demonstrate the generalizability of our pre-trained model, we conduct extensive experiments on various benchmarks with ten downstream tasks. The superior performance of our method is reported in comparison to state-of-the-art pre-training methods, with Dice Score improvement of 0.28\%-14.47\% across six segmentation benchmarks and a consistent accuracy boost of 0.65\%-18.07\% in four individual image classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。