用合成病变训练模型,让CT自动识别冠心病更准更通用
CORA: Generalizable coronary artery disease assessment and risk stratification from coronary CT angiography using pathology-centric representation learning
- 用生成引擎在无标注数据中插入虚拟斑块,聚焦病灶特征学习
- 在9家医院数据上表现优于主流自监督方法,跨中心泛化性强
- 结合临床数据可预测心脏病风险,适合临床部署
冠状动脉疾病是全球主要心血管死亡原因,可通过冠状动脉计算机断层扫描血管成像(CCTA)非侵入式评估。尽管深度学习推动了自动化分析,但受限于专家标注数据稀缺和病灶空间稀疏性(仅占扫描极小部分)。现有无标签预训练策略如掩码图像建模、对比学习侧重全局解剖重建,难以捕捉微小局部病灶。本文提出CORA模型,采用基于合成的自监督策略:通过解剖引导引擎在未标注扫描中插入多样化的钙化与非钙化病变,将预训练重构为异常检测任务,使表征学习偏向临床相关疾病特征。模型在10,138个未标注CCTA体积上预训练,并在九家独立医院数据集上评估。在斑块表征、狭窄检测和冠脉分割任务中,CORA持续优于强基线自监督方法,尤其在外部多中心数据上提升显著,表明其在分布偏移下的鲁棒泛化能力。将影像编码器与结构化临床变量结合,进一步实现近中期主要不良心脏事件(MACE)风险分层。结果表明,以病变为中心、基于合成的预训练是高效、可扩展的冠心病评估策略。
原文摘要 · Abstract (English)
Coronary artery disease, a leading cause of cardiovascular mortality worldwide, can be assessed non-invasively by coronary computed tomography angiography (CCTA). Although deep learning has advanced automated CCTA analysis, clinical translation remains constrained by the scarcity of expert-annotated data and by the spatial sparsity of coronary pathology, which occupies only a small fraction of each scan. Widely used label-free pretraining strategies, such as masked image modeling and contrastive learning, optimize for global anatomical reconstruction and tend to under-represent these tiny localized pathological features. Here we present CORA, an annotation-efficient model for comprehensive coronary artery disease assessment. Rather than reconstructing background anatomy, CORA learns from volumetric CCTA through a synthesis-driven self-supervised strategy: an anatomy-guided engine inserts diverse synthetic calcified and non-calcified lesions into unlabeled scans, reframing pretraining as an abnormality-detection task that biases representation learning toward clinically relevant disease features. We pretrained CORA on 10,138 unlabeled CCTA volumes and evaluated it across datasets from nine independent hospitals. Across plaque characterization, stenosis detection, and coronary artery segmentation, CORA consistently outperformed strong self-supervised pretraining baselines, with the largest gains on external multi-center data, indicating robust generalization under distributional shift. Coupling the imaging encoder with structured clinical variables further enabled near-term major adverse cardiac event (MACE) risk stratification. Our results show that pathology-centric, synthesis-driven pretraining is an effective and scalable strategy for annotation-efficient coronary artery disease assessment from CCTA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。