构建可统一建模多种脑影像的稀疏基础模型,提升临床与科研场景表现。
Learning Sparse Latent Predictive Foundation Model for Multimodal Neuroimaging

- 采用稀疏专家混合架构结合潜在预测目标,统一编码T1w/T2w/FLAIR三种序列
- 在超150万张扫描数据上预训练,25项临床任务中表现优于现有模型
- 适用于多中心、多模态脑影像研究,尤其适合需要稳定泛化的医疗场景
脑部MRI通常包含多种互补序列,如解剖结构清晰的T1加权(T1w)和对液体敏感的T2加权(T2w)图像,以及抑制液体信号的FLAIR序列。然而,在健康系统规模下学习跨多种成像对比机制的统一表征方法仍不成熟。本文提出Neuro-JEPA,一种基于稀疏专家混合架构与潜在预测目标的多模态神经影像基础模型,用于编码核心的T1w、T2w及FLAIR影像。该模型在经过模态特异性预处理后,使用来自428,647例研究的1,551,862张扫描数据进行预训练。我们在三个医疗机构(纽约大学朗格尼、长岛分校、麻省总医院)的25项临床任务,以及12个公开数据集的22项任务中评估了其性能,涵盖单模态、多模态与跨域设置。结果显示,现有神经影像基础模型在多数任务中仅小幅优于简单卷积网络(CNN)基线,而Neuro-JEPA在所有设置中均表现出更强且更一致的性能,验证了其可扩展的方法框架,并强调需引入简单基线、临床异质人群和受控多模态对比来完善基础模型评估体系。
原文摘要 · Abstract (English)
Brain MRIs are routinely acquired as multiple complementary sequences with unique contrast weighting, including T1-weighed imaging (T1w) anatomic and fluid-sensitive T2-weighted (T2w) contrasts. However, methods for learning unified representations across the multitude of MRI contrast mechanisms at health-system scale are lacking. In this study, we introduce Neuro-JEPA, a sparse multimodal neuroimaging foundation model that combines a latent predictive objective with a Mixture-of-Experts architecture to encode brain MRI across core T1w, T2w, and fluid-suppressed FLAIR imaging (FLAIR). We further provide a systematic methodological study of architectural, masking, objective, and sparsity design choices beneficial for robust neuroimaging multimodal representation learning. Neuro-JEPA was pretrained on 1,551,862 scans from 428,647 studies after modality-specific preprocessing with data curation across three core structural brain MRI sequences. We evaluated the learned representations across clinical and research settings, including 25 tasks from three health systems: NYU Langone, NYU Long Island, and Massachusetts General Hospital, and 22 tasks from 12 public datasets, covering unimodal, multimodal and cross-domain evaluation configurations. Across these benchmarks, existing neuroimaging foundation models showed inconsistent gains over a simple convolutional neural network (CNN) baseline, whereas Neuro-JEPA achieved stronger and more consistent performance across all evaluated settings. These results establish a scalable methodological framework for multimodal neuroimaging representation learning and highlight the need for foundation model evaluation protocols that include simple baselines, clinically heterogeneous cohorts and controlled multimodal comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。