提升医学影像诊断的可靠性,防止模型在模糊情况下的过度自信。
Adaptive Conformal Prediction for Reliable and Explainable Medical Image Classification

- 提出自适应λ准则,优化最差情况下的覆盖率,避免关键样本漏判。
- 在器官AMNIST数据集上实现95.72%全局覆盖率,平均预测集大小仅1.09。
- 结果可解释性强,多标签预测与解剖模糊区域注意力高度相关。
医学影像深度学习模型常表现出过度自信,在诊断不确定场景中带来安全隐患。虽然共形预测(CP)提供无分布假设的统计保证,但标准方法如正则化自适应预测集(RAPS)以平均效率为目标,可能掩盖复杂输入上的严重失效。本文提出一种自适应λ准则的RAPS,最小化不同预测集大小分层中的最坏情况覆盖偏差。在包含58,850张腹部CT图像、11个类别的OrganAMNIST数据集上,标准大小优化的RAPS趋近确定性行为,且在不确定样本上出现分层覆盖不足;而本文方法实现95.72%全局覆盖率,平均预测集大小为1.09,并确保所有分层覆盖率不低于90%。在包含107,180张病理图像、9个类别的PathMNIST上交叉验证确认其泛化能力。定量Grad-CAM分析显示(rho = -0.30,p < 1e-22),多标签预测对应于对解剖模糊区域的集中关注。结果表明,该方法在保持效率的同时显著提升可靠性,适用于高安全要求的医疗AI应用。
原文摘要 · Abstract (English)
Deep learning models for medical imaging often exhibit overconfidence, creating safety risks in ambiguous diagnostic scenarios. While Conformal Prediction (CP) provides distribution-free statistical guarantees, standard methods such as Regularized Adaptive Prediction Sets (RAPS) optimize for average efficiency and can mask severe failures on difficult inputs. We propose an Adaptive Lambda Criterion for RAPS that minimizes the worst-case coverage violation across prediction set size strata. On OrganAMNIST (58,850 abdominal CT images, 11 classes), standard size-optimized RAPS converges to near-deterministic behavior with stratified undercoverage on uncertain samples, while our method achieves 95.72 percent global coverage with average set size 1.09 and at least 90 percent coverage across all strata. Cross-domain validation on PathMNIST (107,180 pathology images, 9 classes) confirms generalizability. Quantitative Grad-CAM analysis (rho = -0.30, p < 1e-22) shows that multi-label predictions correspond to focused attention on anatomically ambiguous regions. These results demonstrate that the proposed method improves reliability while maintaining efficiency, making it suitable for safety-critical medical AI applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。