首个专为3D MRI设计的视觉基础模型,提升医学影像分割与分类性能。
Triad: Vision Foundation Model for 3D Magnetic Resonance Imaging
- 基于13万张3D MRI构建自编码器预训练模型,学习通用特征表示
- 在17个分割任务中平均提升2.51%,分类与配准任务分别提升3.97%和4.00%
- 适合从事MRI影像分析的研究者,尤其关注多器官、跨域应用的场景
视觉基础模型(VFMs)通过大规模图像数据预训练,学习通用表征并适用于多种下游任务。然而,现有模型主要基于3D CT数据预训练,而CT与磁共振成像(MRI)在成像原理、信号特性及数据分布上存在显著差异,限制了其在MRI任务中的表现。本文提出面向3D MRI的视觉基础模型Triad,采用广泛使用的自编码器架构,在131,170个3D MRI体积数据上进行预训练,使用器官无关的图像描述约束视觉模态语义分布。该预训练数据集命名为Triad-131K,目前是最大的3D MRI预训练数据集。我们在两个数据模态(同域与跨域)设置下,利用25个下游数据集评估了Triad在器官/肿瘤分割、器官/癌症分类和医学图像配准三个任务上的表现。以Triad预训练权重初始化模型后,nnUNet-Triad在17个数据集上的分割性能相比nnUNet-Scratch提升2.51%;Swin-B-Triad在5个分类任务中较Swin-B-Scratch提升3.97%;SwinUNETR-Triad在2个配准任务中较SwinUNETR-Scratch提升4.00%。研究结果表明,当上游与下游任务的数据模态和器官一致时,预训练可有效提升性能。
原文摘要 · Abstract (English)
Vision foundation models (VFMs) are pre-trained on extensive image datasets to learn general representations for diverse types of data. These models can subsequently be fine-tuned for specific downstream tasks, significantly boosting performance across a broad range of applications. However, existing vision foundation models that claim to be applicable to various clinical tasks are mostly pre-trained on 3D computed tomography (CT), which benefits from the availability of extensive 3D CT databases. Significant differences between CT and magnetic resonance imaging (MRI) in imaging principles, signal characteristics, and data distribution may hinder their practical performance and versatility in MRI-specific applications. Here, we propose Triad, a vision foundation model for 3D MRI. Triad adopts a widely used autoencoder architecture to learn robust representations from 131,170 3D MRI volumes and uses organ-independent imaging descriptions to constrain the semantic distribution of the visual modality. The above pre-training dataset is called Triad-131K, which is currently the largest 3D MRI pre-training dataset. We evaluate Triad across three tasks, namely, organ/tumor segmentation, organ/cancer classification, and medical image registration, in two data modalities (within-domain and out-of-domain) settings using 25 downstream datasets. By initializing models with Triad's pre-trained weights, nnUNet-Triad improves segmentation performance by 2.51% compared to nnUNet-Scratch across 17 datasets. Swin-B-Triad achieves a 3.97% improvement over Swin-B-Scratch in classification tasks across five datasets. SwinUNETR-Triad improves by 4.00% compared to SwinUNETR-Scratch in registration tasks across two datasets. Our study demonstrates that pre-training can improve performance when the data modalities and organs of upstream and downstream tasks are consistent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。