arXiv:2606.13315cs.CVeess.IV2026-06被引 1

提出两种自监督模型,提升脑MRI疾病检测精度。

Masked and Predictive Self-Supervised Foundation Models for 3D Brain MRI

论文配图:Masked and Predictive Self-Supervised Foundation Models for 3D Brain MRI
图 1 · 摘自论文原文
  • 用掩码自编码器和联合嵌入预测架构做自监督预训练。
  • 频域重建损失使模型更敏感于细微解剖结构,提升检测效果。
  • 适合关注医学影像自监督学习与脑病检测的研究者。

自监督基础模型在医学影像中展现出巨大潜力。然而,现有研究多聚焦分割与密集预测任务,对基于MRI的疾病检测自监督基础模型的系统性研究仍较缺乏。本文探讨了两种主要自监督预训练范式:基于重建的掩码自编码器(MAE)和基于预测的联合嵌入预测架构(JEPA)。通过引入新颖的频域重建损失增强MAE对精细解剖结构的敏感性,并在JEPA框架中集成方差-协方差正则化(VCR),促使潜在表示去相关。模型在异构单对比度MRI数据上进行无模态拼接的对比无关预训练。在五个下游疾病检测任务中,结果表明自监督目标设计对医学基础模型预训练至关重要,其下游收益取决于目标与任务结构的相关性:当判别信号由强高频解剖结构主导时,频域正则化表现最佳;当判别信息分布在多个去相关特征维度时,协方差正则化最有效。带有频域监督的MAE在所有任务中均取得最优性能。这表明医学影像中的自监督目标蕴含特定偏差,其下游效益根本上依赖于任务结构。

原文摘要 · Abstract (English)

Self-supervised foundation models have shown strong promise in medical imaging. However, existing MRI foundation-model studies have primarily emphasized segmentation and dense prediction tasks, while systematic investigation of self-supervised foundation models for MRI-based disease detection remains limited. In this work, we investigate two major self-supervised pretraining paradigms for MRI-based disease detection: reconstruction-based learning via Masked Autoencoders (MAE) and predictive representation learning via Joint Embedding Predictive Architectures (JEPA). We study the role of auxiliary objectives by introducing a novel spectral-domain reconstruction loss for MAE to enhance sensitivity to fine-grained anatomical structure, and by integrating variance--covariance regularization (VCR) within our JEPA framework to encourage decorrelated latent representations. Our models are pretrained on heterogeneous single-contrast MRI volumes in a contrast-agnostic setting, without modality concatenation. Across five downstream disease detection tasks, our results highlight the importance of self-supervised objective design for medical foundation model pretraining, demonstrating that the downstream benefit of each objective is determined by its relevance to the task's structure. Specifically, spectral regularization yields the largest improvements when the downstream discriminative signal is characterized by strong high-frequency anatomical structures, while covariance regularization is most beneficial when discriminative information spans multiple decorrelated feature dimensions. MAE with spectral-domain supervision consistently achieves superior downstream performance for MRI-based disease detection. These findings suggest that self-supervised objectives in medical imaging encode specific biases, and their downstream benefit is fundamentally conditioned on the task's structure.

自监督学习脑MRI疾病检测MAE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。