构建冠脉造影像素级分类新基准,助力心脏病智能诊断研究
CARDIAG: A Dense Segment Classification Benchmark of Deep Learning Architectures for Coronary Angiography

- 设计多中心多标签数据集,支持24种模型密集分类评估
- 最优模型达宏F1=0.479,融合结构提升性能并验证校准性
- 揭示高/低分辨率特征对建模的重要性,适合医学影像研究者
精准的冠脉造影像素级分类对心血管疾病评估至关重要,但领域内缺乏标准化评估协议。本文提出CARDIAG——一个用于评估深度学习模型在冠脉造影中像素级分类的新型基准,将像素划分为SYNTAX类别或背景。评估涵盖24种架构,从经典卷积网络到近期基于状态空间的视觉算法。我们发布了CARDIAG数据集,包含多中心、多标签数据,并精心划分以可靠计算指标,考虑直径误差、重叠度、中心线质量及校准性。数据含SYNTAX标签、二值掩码、置信度图、分割掩码及中间帧,以及选定的非敏感DICOM元数据。实验表明,使用ConvNeXt V2编码器与DeepLab V3 Plus解码器的模型表现最佳,宏F1达0.456,再与Mamba U-Net和特征金字塔网络集成后,F1提升至0.479。所有模型均表现出良好校准性,我们进一步分析了前五名方法的泛化能力与数据效率。结果强调了高分辨率与低分辨率特征在编码中的共同重要性,并验证了模型在患者人口统计、血管侧别与投影视角配置下的正确性。该基准为未来研究提供了稳健严谨的评估基础,不仅适用于SYNTAX分割,还可拓展至病灶检测等任务。
原文摘要 · Abstract (English)
Accurate pixel-level classification of coronary angiograms is critical for cardiovascular disease assessment, yet the field lacks standardized evaluation protocols. In this work we demonstrate a new benchmark for the assessment of deep learning models which densely classify pixels of coronary angiograms to one of SYNTAX classes (or background). The evaluation covers 24 distinct architectures starting with classic convnets to recent state-space-based vision algorithms. We release CARDIAG - a multi-center, multi-label dataset which we carefully split to reliably compute metrics, accounting for diameter error, overlap, centerline quality and calibration. The data contains SYNTAX labels, binary, uncertainty and segmentation masks as well as intermediate frames together with the selected non-sensitive DICOM metadata. From the multitude of algorithms, we nominate ConvNeXt V2 encoder with DeepLab V3 Plus decoder as the best performing, achieving macro $F_1=0.456$, which we then ensemble with Mamba U-Net and Feature Pyramid Network, for an increased $F_1=0.479$. We demonstrate all the architectures to be well calibrated and determine the generalization of the top 5 methods, together with the data efficiency of these architectures. We highlight the importance of both high-resolution and low-resolution features in encoding. We also demonstrate the model correctness in the context of patient demographic, vessel sides and projection angle configurations. Overall the released benchmark allows for future studies to robustly and rigorously assess the proposals, not only for SYNTAX segmentation, but lesion detection and many more.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。