arXiv:2507.22024eess.IVcs.CV2025-07被引 1

Cardiac-CLIP让3D心脏CT图像理解更智能,支持临床关键病预测。

Cardiac-CLIP: A Vision-Language Foundation Model for 3D Cardiac CT Images

  • 分两阶段预训练:先用3D掩码自编码器学解剖特征,再用对比学习对齐图文表示。
  • 在1.3万例真实临床数据上训练,跨机构测试表现领先现有方法。
  • 可精准辅助诊断急性冠脉综合征等复杂临床任务,适合医学AI研究者参考。

基础模型在医疗领域展现出巨大潜力,但在复杂心血管诊断中的应用仍不充分。本文提出Cardiac-CLIP,一种面向3D心脏CT图像的多模态基础模型。该模型采用两阶段预训练策略:第一阶段使用3D掩码自编码器(MAE)对大规模无标签体数据进行自监督表征学习,使视觉编码器捕捉丰富的解剖与上下文特征;第二阶段引入对比学习对齐视觉与文本表示,实现跨模态理解。为支持预训练,我们收集了16,641例真实临床CT扫描,并补充11.4万份公开数据。同时,将自由文本放射科报告标准化为统一模板,基于诊断属性构建病理向量,生成软标签矩阵以监督对比学习过程。为全面评估模型性能,我们从12家独立医疗机构收集6,722例真实临床数据,并结合开源数据构建评测集。Cardiac-CLIP在心血管异常分类、信息检索和临床分析等多个任务中进行评估,实验表明其在内部及外部数据上均达到当前最优表现,尤其在预测急性冠脉综合征这一现实场景下表现出色。

原文摘要 · Abstract (English)

Foundation models have demonstrated remarkable potential in medical domain. However, their application to complex cardiovascular diagnostics remains underexplored. In this paper, we present Cardiac-CLIP, a multi-modal foundation model designed for 3D cardiac CT images. Cardiac-CLIP is developed through a two-stage pre-training strategy. The first stage employs a 3D masked autoencoder (MAE) to perform self-supervised representation learning from large-scale unlabeled volumetric data, enabling the visual encoder to capture rich anatomical and contextual features. In the second stage, contrastive learning is introduced to align visual and textual representations, facilitating cross-modal understanding. To support the pre-training, we collect 16641 real clinical CT scans, supplemented by 114k publicly available data. Meanwhile, we standardize free-text radiology reports into unified templates and construct the pathology vectors according to diagnostic attributes, based on which the soft-label matrix is generated to supervise the contrastive learning process. On the other hand, to comprehensively evaluate the effectiveness of Cardiac-CLIP, we collect 6,722 real-clinical data from 12 independent institutions, along with the open-source data to construct the evaluation dataset. Specifically, Cardiac-CLIP is comprehensively evaluated across multiple tasks, including cardiovascular abnormality classification, information retrieval and clinical analysis. Experimental results demonstrate that Cardiac-CLIP achieves state-of-the-art performance across various downstream tasks in both internal and external data. Particularly, Cardiac-CLIP exhibits great effectiveness in supporting complex clinical tasks such as the prospective prediction of acute coronary syndrome, which is notoriously difficult in real-world scenarios.

医学影像多模态3D建模心脏疾病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。