系统评估医学影像AI性能置信区间的可靠性,给出实用决策指南。
Performance uncertainty in medical image analysis: a large-scale investigation of confidence intervals
- 在24个任务上测试19种模型,全面分析不同置信区间方法的表现。
- 置信区间精度受样本量、评估指标和聚合策略影响,差异可达数倍。
- 提出决策树工具,帮助研究者选择适合的置信区间方法。
性能不确定性量化对医学影像人工智能(AI)的可靠验证和临床转化至关重要。置信区间(CIs)在此过程中起核心作用,用于反映性能估计的精确性。然而,由于缺乏针对医学影像中置信区间行为的大规模研究,学界对现有多种置信区间方法及其在特定场景下的表现仍不了解。本研究旨在填补这一空白。我们对24个分割与分类任务进行了大规模实证分析,每个任务使用19个训练好的模型,涵盖广泛常用性能指标、多种聚合策略及多个主流置信区间方法。在所有设置下评估每种方法的可靠性(覆盖概率)和精度(区间宽度),以揭示其对研究特征的依赖关系。研究发现:1)可靠置信区间的样本量需求从几十到数千例不等,视具体参数而定;2)置信区间表现受性能指标选择显著影响;3)聚合策略显著影响置信区间可靠性,例如宏观平均需要更多样本;4)机器学习任务类型(分割与分类)调节上述效应;5)不同置信区间方法在不同应用场景下的可靠性和精度不均等。最后,我们基于结果构建了实用决策树,为社区提供指导,推动未来性能不确定性报告的共识规范制定。
原文摘要 · Abstract (English)
Performance uncertainty quantification is essential for reliable validation and eventual clinical translation of medical imaging artificial intelligence (AI). Confidence intervals (CIs) play a central role in this process by indicating how precise a reported performance estimate is. Yet, due to the limited amount of work examining CI behavior in medical imaging, the community remains largely unaware of how many diverse CI methods exist and how they behave in specific settings. The purpose of this study is to close this gap. To this end, we conducted a large-scale empirical analysis across a total of 24 segmentation and classification tasks, using 19 trained models per task group, a broad spectrum of commonly used performance metrics, multiple aggregation strategies, and several widely adopted CI methods. Reliability (coverage) and precision (width) of each CI method were estimated across all settings to characterize their dependence on study characteristics. Our analysis revealed five principal findings: 1) the sample size required for reliable CIs varies from a few dozens to several thousands of cases depending on study parameters; 2) CI behavior is strongly affected by the choice of performance metric; 3) aggregation strategy substantially influences the reliability of CIs, e.g. they require more observations for macro than for micro; 4) the machine learning problem (segmentation versus classification) modulates these effects; 5) different CI methods are not equally reliable and precise depending on the use case. Finally, we derived practical implications of this study in the form of a decision tree which shall prove useful to the community. This paves the way for future consensus guidelines on reporting performance uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。