基于视频文本对比学习的冠脉造影多视角分析模型,实现精准自动解读。
DeepCORO-CLIP: A Multi-View Foundation Model for Comprehensive Coronary Angiography Video-Text Analysis and External Validation
- 通过多视角视频-文本对比学习,融合多个投影视角进行整体评估。
- 显著狭窄检测AUROC达0.89,定量分析误差低于临床报告(13.6% vs 19.0%)。
- 可支持心梗风险预测、心功能估算及疾病进展追踪,适合临床部署。
冠状动脉造影是评估冠心病的金标准,但视觉解读存在读者间差异。现有人工智能方法多聚焦单帧或单投影的狭窄检测,难以全面评估。本文提出DeepCORO-CLIP,一个在蒙特利尔心脏研究所28,117名患者共32,473项研究的203,808段冠脉造影视频上训练的多视图基础模型,经旧金山大学4,249项研究外部验证。该模型采用注意力池化整合多投影信息,实现诊断、预后及疾病进展任务的全片级评估。显著狭窄检测内部AUROC为0.888,外部验证达0.89;与核心实验室定量冠脉造影相比,平均绝对误差为13.6%,优于临床报告的19.0%。在慢性完全闭塞、血管内血栓、钙化检测上表现优异。迁移学习使一年主要心血管事件预测的AUROC达0.79,左室射血分数估计均值绝对误差为7.3%。嵌入向量还可捕捉序列检查中的疾病进展。医院部署平均推理时间仅4.2秒,代码、数据、模型权重和部署架构已公开。
原文摘要 · Abstract (English)
Coronary angiography is the reference standard for evaluating coronary artery disease, yet visual interpretation remains variable between readers. Existing artificial intelligence methods typically analyze single frames or projections and focus mainly on stenosis, limiting comprehensive coronary assessment. We present DeepCORO-CLIP, a multi-view foundation model trained with video-text contrastive learning on 203,808 angiography videos from 28,117 patients across 32,473 studies at the Montreal Heart Institute and externally validated on 4,249 studies from the University of California, San Francisco. DeepCORO-CLIP integrates multiple projections with attention-based pooling for study-level assessment across diagnostic, prognostic, and disease progression tasks. For significant stenosis detection, the model achieved an AUROC of 0.888 internally and 0.89 on external validation. Mean absolute error against core laboratory quantitative coronary angiography was 13.6%, lower than clinical reports at 19.0%. The model also performed strongly for chronic total occlusion, intracoronary thrombus, and coronary calcification detection. Transfer learning enabled prediction of one-year major adverse cardiovascular events with AUROC 0.79 and estimation of left ventricular ejection fraction with mean absolute error 7.3%. Embeddings also captured disease progression across serial examinations. With a mean inference time of 4.2 seconds in hospital deployment, DeepCORO-CLIP provides a foundation for automated coronary angiography interpretation at the point of care. Code, sample data, model weights, and deployment infrastructure are publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。