提升医学影像AI的可信度,让模型自知何时不准。
Confidence-Uncertainty Boundary Calibration for Bayesian Deep Learning in Medical Image Analysis
- 用新损失函数让模型高信心时不出错,低信心时别漏诊。
- 在肺炎、糖尿病眼病等3个任务上显著改善不确定性校准。
- 适合临床部署,尤其数据少或类别不平衡场景下更可靠。
在基于医学影像的临床决策支持系统中,AI判断的可靠性与预测准确率同样重要。尽管深度学习模型已展现高精度,但常出现错误预测时过于自信的问题。为促进临床采纳,模型需能以与预测正确性相关的方式量化不确定性,使医生可识别不可靠输出进行复核。本文提出基于贝叶斯深度学习的概率优化框架,首次引入置信度-不确定性边界曲线(CUBC)作为中间目标。基于此目标,设计新型置信度-不确定性边界损失(CUB-Loss),在训练中正则化置信度与不确定性估计的一致性,对高置信度错误和低置信度正确预测施加惩罚。训练完成后,引入边界曲线校准误差(BCCE)度量模型边界对齐程度,并提出双温度缩放(DTS)策略进行后处理校准,进一步调整不同置信度-不确定性区域的后验分布。该框架在三种医学影像任务中验证:肺炎自动筛查、糖尿病视网膜病变检测及皮肤病变识别。实证结果表明,该方法在多种模态下显著提升不确定性校准效果,在数据稀缺和严重不平衡数据集上仍保持稳健性能,具备实际临床应用潜力。
原文摘要 · Abstract (English)
In critical decision support systems based on medical imaging, the reliability of AI-assisted decision-making is as relevant as predictive accuracy. Although deep learning models have demonstrated significant accuracy, they frequently suffer from miscalibration, manifested as overconfidence in erroneous predictions. To facilitate clinical acceptance, it is imperative that models quantify uncertainty in a manner that correlates with prediction correctness, allowing clinicians to identify unreliable outputs for further review. To address this necessity, this paper proposes a probabilistic optimization framework grounded in Bayesian deep learning. Specifically, the Confidence-Uncertainty Boundary Curve (CUBC) is first explored as an intermediate operational target. Grounded in this target, a novel Confidence-Uncertainty Boundary Loss (CUB-Loss) is proposed to regularize the alignment between prediction confidence and uncertainty estimates during training, imposing penalties on high-certainty errors and low-certainty correct predictions. Upon completion of training optimization, a Boundary Curve Calibration Error (BCCE) metric is further introduced to measure the degree of boundary alignment in the calibrated model. Building on this measurement, a Dual Temperature Scaling (DTS) strategy is devised to perform post-hoc refinement, further adjusting the posterior predictive distribution across different confidence-uncertainty regions. The proposed framework is validated on three distinct medical imaging tasks: automatic screening of pneumonia, diabetic retinopathy detection, and identification of skin lesions. Empirical results demonstrate that the proposed approach improves uncertainty calibration across diverse modalities, maintains robust performance in data-scarce scenarios, and remains effective on severely imbalanced datasets, underscoring its potential for real clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。