自适应掩码自动编码器提升多对比度3D医学影像预训练效果
Self Pre-training with Adaptive Mask Autoencoders for Variable-Contrast 3D Medical Imaging
- 设计可处理不同数量输入对比度的3D自适应掩码自动编码器
- 在脑梗死分割任务上提升2.8%-3.7%性能
- 适用于多模态医学影像分析,尤其适合真实MRI数据
掩码自动编码器(MAE)在自然图像视觉变换器(ViT)预训练中表现出色。通过从部分遮蔽的输入重建完整图像,ViT编码器能聚合上下文信息以预测缺失区域。这一能力在医学影像中尤为重要,因为解剖结构与其周围区域存在功能和力学关联。然而,现有方法未考虑实际磁共振(MR)研究中输入图像数量的变化。为此,我们提出3D自适应掩码自动编码器(AMAE)架构,可适应每名受试者不同数量的3D对比度输入。使用包含45,364名受试者的MRI数据集进行预训练,1,648名用于训练、193名用于验证、215名用于测试。结果表明,该自适应掩码自动编码器的自监督预训练可使基于ViT的分割模型在脑梗死分割任务上提升2.8%-3.7%。
原文摘要 · Abstract (English)
The Masked Autoencoder (MAE) has recently demonstrated effectiveness in pre-training Vision Transformers (ViT) for analyzing natural images. By reconstructing complete images from partially masked inputs, the ViT encoder gathers contextual information to predict the missing regions. This capability to aggregate context is especially important in medical imaging, where anatomical structures are functionally and mechanically linked to surrounding regions. However, current methods do not consider variations in the number of input images, which is typically the case in real-world Magnetic Resonance (MR) studies. To address this limitation, we propose a 3D Adaptive Masked Autoencoders (AMAE) architecture that accommodates a variable number of 3D input contrasts per subject. A magnetic resonance imaging (MRI) dataset of 45,364 subjects was used for pretraining and a subset of 1648 training, 193 validation and 215 test subjects were used for finetuning. The performance demonstrates that self pre-training of this adaptive masked autoencoders can enhance the infarct segmentation performance by 2.8%-3.7% for ViT-based segmentation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。