arXiv:2608.30975cs.CVcs.AI2026-08中稿 · STACOM 2026

基于自监督学习的通用心脏MRI视频模型,可统一处理多种序列数据。

MR-JEPA: A General Purpose Video Foundation Model for Cardiac MRI

  • 通过管状令牌化与时空掩码增强,扩展JEPA至3D时空输入。
  • 在6个下游任务中均超越基线,应变任务误差降低21%-27%。
  • 无需标注即可预训练,适合临床心脏定量与诊断应用。

心脏磁共振成像(CMR)产生丰富的时序数据,如动态电影视频和空间LGE/映射堆栈,但多数深度学习方法仅处理单个2D切片,忽略上下文信息。本文提出MR-JEPA,一种针对CMR的自监督视频基础模型,通过管状令牌化、时空掩码增强及从2D CMR基础模型初始化,将LeJEPA扩展至3D时空输入。不同于以往仅限于电影数据的CMR视频模型,MR-JEPA在来自两个中心的10,505例患者多序列数据(电影、LGE、映射)上进行无标注预训练。我们使用统一的多视图门控注意力架构,在六个下游任务上评估冻结编码器:左室射血分数(LV EF)、右室射血分数(RV EF)、三种心肌应变(GLS、GCS、GRS)及四分类疾病检测。MR-JEPA在所有五个回归任务中均优于对比方法,包括一个在更多数据上用文本监督预训练的领域特定模型和一个自然视频基础模型,实现LV EF MAE为4.79%(r=0.764),GLS MAE为1.87(r=0.805),应变任务误差降低21%-27%。疾病检测任务中,宏平均AUC达0.868,虽采用完全自监督预训练目标,仍保持与领域特定基线相当的性能。结果表明,统一视频编码器在临床心脏量化与诊断中具有潜力,可高效利用多样化CMR序列。

原文摘要 · Abstract (English)

Cardiac magnetic resonance imaging (CMR) produces rich sequential data such as temporal cine videos and spatial LGE/mapping stacks, yet most deep learning approaches process individual 2D slices, discarding this context. We present MR-JEPA, a self-supervised video foundation model for CMR that extends LeJEPA to 3D spatiotemporal inputs through tubelet tokenization, spatiotemporal masking augmentation, and initialization from a 2D CMR foundation model. Unlike prior CMR video models limited to cine data, MR-JEPA is pretrained on multi-sequence data (cine, LGE, mapping) from 10,505 patients across two centers without annotations. We evaluate the frozen encoder on six downstream tasks using a unified multi-view gated attention architecture: LV ejection fraction, RV ejection fraction, three myocardial strains (GLS, GCS, GRS), and four-class disease detection. MR-JEPA outperforms other compared methods on all five regression tasks, including both a domain-specific CMR model pretrained on more data with text supervision and a natural-video foundation model, achieving an LV EF MAE of 4.79% (r =0.764) and a GLS MAE of 1.87 (r=0.805), with 21-27% MAE reductions over baselines on strain tasks. For disease detection, MR-JEPA achieved a macro AUG of 0.868, remaining competitive with the domain-specific baseline despite using a fully self-supervised pretraining objective. These results demonstrate the potential of a unified video encoder for robust, multi-view utilization of diverse CMR sequences in clinical cardiac quantification and diagnosis.

心脏MRI自监督学习视频基础模型医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。