arXiv:2505.08845eess.IVcs.AI2025-05

用专家标注验证宫颈非典型性分类中的置信预测效果

Validation of Conformal Prediction in Cervical Atypia Classification

  • 引入专家标注数据评估置信预测集的准确性与实用性
  • 发现传统覆盖率评估高估性能,预测集常包含不合理的类
  • 揭示置信预测在识别模糊和分布外样本上的潜力

基于深度学习的宫颈癌分类可提升资源匮乏地区的筛查可及性。然而,深度学习模型常过度自信,无法可靠反映诊断不确定性。通常优化为最大似然预测,难以传达结果中的不确定或模糊性。这些问题可通过置信预测(conformal prediction)框架解决,该框架可生成包含可能类别的预测集,其大小反映模型不确定性,信心越高则集合越小。但现有评估多关注预测集是否包含真实类别,忽略其中可能包含的额外错误类别。我们主张预测集应真实且对使用者有价值,所列可能类别需符合人类预期,而非过于宽松地包含假阳性或不合理类别。本研究使用多位标注者收集的专家标注数据,全面验证三种置信预测方法在三个宫颈非典型性分类深度学习模型上的表现。分析显示,基于覆盖率的传统评估过高估计性能,当前方法生成的预测集常与人类标注不一致。此外,我们探讨了置信预测在识别模糊和分布外数据方面的能力。

原文摘要 · Abstract (English)

Deep learning based cervical cancer classification can potentially increase access to screening in low-resource regions. However, deep learning models are often overconfident and do not reliably reflect diagnostic uncertainty. Moreover, they are typically optimized to generate maximum-likelihood predictions, which fail to convey uncertainty or ambiguity in their results. Such challenges can be addressed using conformal prediction, a model-agnostic framework for generating prediction sets that contain likely classes for trained deep-learning models. The size of these prediction sets indicates model uncertainty, contracting as model confidence increases. However, existing conformal prediction evaluation primarily focuses on whether the prediction set includes or covers the true class, often overlooking the presence of extraneous classes. We argue that prediction sets should be truthful and valuable to end users, ensuring that the listed likely classes align with human expectations rather than being overly relaxed and including false positives or unlikely classes. In this study, we comprehensively validate conformal prediction sets using expert annotation sets collected from multiple annotators. We evaluate three conformal prediction approaches applied to three deep-learning models trained for cervical atypia classification. Our expert annotation-based analysis reveals that conventional coverage-based evaluations overestimate performance and that current conformal prediction methods often produce prediction sets that are not well aligned with human labels. Additionally, we explore the capabilities of the conformal prediction methods in identifying ambiguous and out-of-distribution data.

宫颈癌置信预测深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。