arXiv:2606.28991cs.CVeess.IV2026-06

利用扫描元数据提升心脏MRI模型表征能力,效果优于纯图像预训练。

Learning from Acquisition: Metadata-driven Multimodal Pre-training for Cardiac MRI

论文配图:Learning from Acquisition: Metadata-driven Multimodal Pre-training for Cardiac MRI
图 1 · 摘自论文原文
  • 将扫描参数转为文本监督信号,实现多模态对比学习
  • 在多种任务上准确率超86%,分割任务表现最优
  • 仅需不到1%的图像量,适合资源受限场景

心脏磁共振成像(CMR)通常记录结构化采集元数据,但多数CMR基础模型依赖纯图像预训练,未充分挖掘这一天然弱语义监督信息。本文提出MetaCLIP-CMR,基于对比语言-图像预训练(CLIP)框架,将成像模态、解剖视角、扫描仪厂商、场强及扫描仪型号转化为文本监督信号,用于CMR表征学习。预训练图像编码器在成像模态分类、电影视图分类和心脏分割任务上评估:分别达到86.8%和86.5%的准确率,显著优于ImageNet与掩码重建初始化。在下游心脏分割任务中,无论全量数据或仅20%微调,其在ACDC与M&Ms cine短轴(SAX)设置下均取得最高Dice分数。相比近期以图像为主的大规模CMR预训练模型,MetaCLIP-CMR在性能相当的前提下,所需预训练图像量不足其1%。结果表明,元数据学习为构建基础级CMR表征提供了自然且高效的监督策略,凸显元数据驱动多模态预训练的潜力。

原文摘要 · Abstract (English)

Cardiac magnetic resonance imaging (CMR) routinely records structured acquisition metadata, yet most CMR foundation models rely primarily on image-only pre-training and leave this naturally available source of weak semantic supervision largely underexplored. We propose MetaCLIP-CMR, a metadata-driven framework based on Contrastive Language--Image Pre-training (CLIP), which converts imaging modality, anatomical view, scanner vendor, field strength, and scanner model into textual supervision for CMR representation learning. The pretrained image encoder is evaluated on imaging modality classification, cine view classification, and cardiac segmentation. MetaCLIP-CMR achieves 86.8% modality accuracy and 86.5% cine view accuracy, clearly outperforming ImageNet and masked reconstruction initialisations. For downstream cardiac segmentation, MetaCLIP-CMR consistently obtains the highest Dice score across the evaluated ACDC and M&Ms cine short-axis (SAX) settings under both full-data and 20% fine-tuning regimes. Compared with recent image-focused large-scale CMR pre-training models, MetaCLIP-CMR achieves comparable ACDC segmentation performance, while requiring less than 1% of their pre-training image scale. These results suggest that metadata learning offers a natural and easy-to-use strategy for transforming routinely recorded acquisition information into effective supervision for foundation-level CMR representation learning, highlighting the promise of metadata-driven multimodal pre-training.

医学影像多模态元数据预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。