用重采样数据评估深度学习心脏影像生物标志物的不确定性与精度
Confidence Matters: Uncertainty Quantification and Precision Assessment of Deep Learning-based CMR Biomarker Estimates Using Scan-rescan Data
- 采用深度集成等技术量化模型预测不确定性
- 点估计精度高但分布重叠不足45%且65%以上差异显著
- 适合关注模型可靠性与临床可重复性的研究者
深度学习在动态心脏磁共振(CMR)分析中的性能通常仅以准确率评估,忽略精确度。本文将深度集成、测试时增强和蒙特卡洛丢弃等不确定性估计方法应用于先进的心脏功能生物标志物估计算法,并提出基于分布的新指标来评估生物标志物的精确度。模型在两个外部验证的扫描-重扫CMR数据集上表现优异:平均Dice系数达87%,点估计精度高。然而,分布指标显示,扫描与重扫置信区间重叠超过50%的情况不足45%;扫描与重扫生物标志物间的统计相似性检验在超过65%的病例中存在显著差异。结论表明,尽管点估计看似良好,但分布分析揭示其精确度较低,强调需采用更全面的指标评估扫描-重扫一致性。
原文摘要 · Abstract (English)
The performance of deep learning (DL) methods for the analysis of cine cardiovascular magnetic resonance (CMR) is typically assessed in terms of accuracy, overlooking precision. In this work, uncertainty estimation techniques, namely deep ensemble, test-time augmentation, and Monte Carlo dropout, are applied to a state-of-the-art DL pipeline for cardiac functional biomarker estimation, and new distribution-based metrics are proposed for the assessment of biomarker precision. The model achieved high accuracy (average Dice 87%) and point estimate precision on two external validation scan-rescan CMR datasets. However, distribution-based metrics showed that the overlap between scan/rescan confidence intervals was >50% in less than 45% of the cases. Statistical similarity tests between scan and rescan biomarkers also resulted in significant differences for over 65% of the cases. We conclude that, while point estimate metrics might suggest good performance, distributional analyses reveal lower precision, highlighting the need to use more representative metrics to assess scan-rescan agreement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。