arXiv:2409.17763cs.CVcs.AI2024-09中稿 · MICCAI 2024 confer…被引 19

医学影像AI论文常忽略性能波动,导致无法判断模型是否适合临床应用。

Confidence intervals uncovered: Are we ready for real-world medical imaging AI?

  • 分析221篇MICCAI 2023分割论文,超50%未评估性能变异,仅1篇报告置信区间。
  • 提出用均值DSC的二次多项式近似标准差,可重建95%置信区间。
  • 超60%论文中第一与第二名模型性能重叠,难以确定优劣,影响临床转化决策。

医学影像正引领医疗AI变革,但性能报告常仅依赖平均值,忽略性能波动。本文分析221篇MICCAI 2023分割论文,发现超过50%未评估性能变异,仅1篇(0.5%)报告了置信区间(CIs)。基于56个前期MICCAI挑战赛外部数据,我们发现模型均值Dice相似系数(DSC)的二次多项式可准确近似标准差(SD),从而重建置信区间。进一步重构2023年论文的95% CIs,中位宽度为0.03,是首尾方法性能差距的三倍。超过60%论文中,第二名模型的均值性能落在第一名的置信区间内。结论:当前发表结果不足以支持模型临床转化决策。

原文摘要 · Abstract (English)

Medical imaging is spearheading the AI transformation of healthcare. Performance reporting is key to determine which methods should be translated into clinical practice. Frequently, broad conclusions are simply derived from mean performance values. In this paper, we argue that this common practice is often a misleading simplification as it ignores performance variability. Our contribution is threefold. (1) Analyzing all MICCAI segmentation papers (n = 221) published in 2023, we first observe that more than 50% of papers do not assess performance variability at all. Moreover, only one (0.5%) paper reported confidence intervals (CIs) for model performance. (2) To address the reporting bottleneck, we show that the unreported standard deviation (SD) in segmentation papers can be approximated by a second-order polynomial function of the mean Dice similarity coefficient (DSC). Based on external validation data from 56 previous MICCAI challenges, we demonstrate that this approximation can accurately reconstruct the CI of a method using information provided in publications. (3) Finally, we reconstructed 95% CIs around the mean DSC of MICCAI 2023 segmentation papers. The median CI width was 0.03 which is three times larger than the median performance gap between the first and second ranked method. For more than 60% of papers, the mean performance of the second-ranked method was within the CI of the first-ranked method. We conclude that current publications typically do not provide sufficient evidence to support which models could potentially be translated into clinical practice.

医学影像性能评估置信区间临床转化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。