arXiv:2603.13590cs.CVcs.AI2026-03

用快速扫描的MRI定位图,结合心电图和病历数据,估算心脏健康指标。

Opportunistic Cardiac Health Assessment: Estimating Phenotypes from Localizer MRI through Multi-Modal Representations

  • 融合定位影像、心电图和患者信息,构建多模态预测框架
  • 对功能型指标预测准确,结构型指标相关性高
  • 仅需定位图即可完成评估,适合临床快速筛查

心血管疾病是主要死因。心脏表型(CPs)如射血分数是评估心脏健康的金标准,但依赖高分辨率动态心脏磁共振(CMR),成本高且耗时。每次磁共振检查前都会生成快速粗糙的定位图用于扫描规划,之后被丢弃。尽管定位图无诊断价值且缺乏时间信息,却能快速提供重要结构信息。同时,患者人口学与生活方式也影响心脏评估。心电图(ECG)成本低、临床常规使用,可捕捉心脏电活动。本文提出C-TRIP(Cardiac Tri-modal Representations for Imaging Phenotypes),一种多模态框架,通过对齐定位影像、ECG信号与表格化元数据,学习鲁棒潜在空间,并仅以定位图为输入预测CPs。该方法利用定位图提供的廉价空间信息、ECG提供的时序信息及表格式患者特征。整体流程分三阶段:先独立训练各模态编码器;再融合预训练编码器统一潜在空间;最后基于融合表示进行表型预测,推理仅依赖定位图。实验表明,所提框架在功能型表型上预测准确,在结构型表型上相关性高。由于定位图本身快速且低成本,该框架有望提升表型评估的可及性。

原文摘要 · Abstract (English)

Cardiovascular diseases are the leading cause of death. Cardiac phenotypes (CPs), e.g., ejection fraction, are the gold standard for assessing cardiac health, but they are derived from cine cardiac magnetic resonance imaging (CMR), which is costly and requires high spatio-temporal resolution. Every magnetic resonance (MR) examination begins with rapid and coarse localizers for scan planning, which are discarded thereafter. Despite non-diagnostic image quality and lack of temporal information, localizers can provide valuable structural information rapidly. In addition to imaging, patient-level information, including demographics and lifestyle, influence the cardiac health assessment. Electrocardiograms (ECGs) are inexpensive, routinely ordered in clinical practice, and capture the temporal activity of the heart. Here, we introduce C-TRIP (Cardiac Tri-modal Representations for Imaging Phenotypes), a multi-modal framework that aligns localizer MRI, ECG signals, and tabular metadata to learn a robust latent space and predict CPs using localizer images as an opportunistic alternative to CMR. By combining these three modalities, we leverage cheap spatial and temporal information from localizers, and ECG, respectively while benefiting from patient-specific context provided by tabular data. Our pipeline consists of three stages. First, encoders are trained independently to learn uni-modal representations. The second stage fuses the pre-trained encoders to unify the latent space. The final stage uses the enriched representation space for CP prediction, with inference performed exclusively on localizer MRI. Proposed C-TRIP yields accurate functional CPs, and high correlations for structural CPs. Since localizers are inherently rapid and low-cost, our C-TRIP framework could enable better accessibility for CP estimation.

心脏评估多模态影像分析低成本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。