提出两种拒答机制,让模型在不确定时主动放弃预测,提升胸片诊断可靠性。
Multi-pathology Chest X-ray Classification with Rejection Mechanisms
- 基于熵和置信区间设计拒答机制,模型可主动回避不确定预测
- 在三个数据集上,熵基拒答使所有病灶平均AUC最高,达0.893
- 适合需要高安全性的临床AI辅助诊断场景,尤其多病共现情况
深度学习模型在高风险医学影像任务中过度自信,尤其在需同时检测多种共存病灶的胸部X光片多标签分类中尤为突出。本研究提出一种基于DenseNet-121骨干网络的不确定性感知框架,引入两种选择性预测机制:基于熵的拒答与基于置信区间拒答。二者使模型可在不确定时主动拒绝预测,将模糊病例交由临床专家处理。采用分位数校准方法,以全局或类别特定策略调整拒答阈值。在三个大型公开数据集(PadChest、NIH ChestX-ray14、MIMIC-CXR)上的实验表明,选择性拒答显著改善了诊断准确率与覆盖率之间的权衡,其中熵基拒答在所有病灶上取得最高平均AUC(0.893)。结果支持将选择性预测集成至AI辅助诊断流程,为临床环境中更安全、具备不确定性意识的深度学习部署提供可行路径。
原文摘要 · Abstract (English)
Overconfidence in deep learning models poses a significant risk in high-stakes medical imaging tasks, particularly in multi-label classification of chest X-rays, where multiple co-occurring pathologies must be detected simultaneously. This study introduces an uncertainty-aware framework for chest X-ray diagnosis based on a DenseNet-121 backbone, enhanced with two selective prediction mechanisms: entropy-based rejection and confidence interval-based rejection. Both methods enable the model to abstain from uncertain predictions, improving reliability by deferring ambiguous cases to clinical experts. A quantile-based calibration procedure is employed to tune rejection thresholds using either global or class-specific strategies. Experiments conducted on three large public datasets (PadChest, NIH ChestX-ray14, and MIMIC-CXR) demonstrate that selective rejection improves the trade-off between diagnostic accuracy and coverage, with entropy-based rejection yielding the highest average AUC across all pathologies. These results support the integration of selective prediction into AI-assisted diagnostic workflows, providing a practical step toward safer, uncertainty-aware deployment of deep learning in clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。