arXiv:2503.11851eess.IVcs.AI2025-03

用双向注意力+不确定性量化,让医疗影像模型更准更可信。

Interpretable Deep Learning Framework for Improved Disease Classification in Medical Imaging

  • 双向交叉注意力融合不同网络特征,增强判别力
  • MC-Dropout与分位数预测结合,输出带置信度的分类结果
  • 在4个数据集上准确率超98%,且能可视化不确定区域

深度学习模型在医学影像分析中应用日益广泛,但常产生过度自信的预测,影响临床可靠性。本文提出一种统一的深度学习框架,提升特征融合、可解释性与预测可靠性。引入跨引导通道空间注意力结构,融合EfficientNetB4与ResNet34提取的特征;双向注意力机制促进不同感受野网络间的信息交互,增强判别性与上下文特征学习。采用蒙特卡洛丢弃(MC-Dropout)结合共形预测进行定量不确定性评估,生成统计有效的预测集,并以熵为基础实现不确定性可视化。在四个医学影像基准数据集上验证:新冠胸片AUC达99.75%,结核病100%,肺炎胸片99.3%,视网膜OCT图像98.69%。不确定性感知推理得到校准的预测集,且可解释不确定性示例,显著提升透明度与可信度。

原文摘要 · Abstract (English)

Deep learning models have gained increasing adoption in medical image analysis. However, these models often produce overconfident predictions, which can compromise clinical accuracy and reliability. Bridging the gap between high-performance and awareness of uncertainty remains a crucial challenge in biomedical imaging applications. This study focuses on developing a unified deep learning framework for enhancing feature integration, interpretability, and reliability in prediction. We introduced a cross-guided channel spatial attention architecture that fuses feature representations extracted from EfficientNetB4 and ResNet34. Bidirectional attention approach enables the exchange of information across networks with differing receptive fields, enhancing discriminative and contextual feature learning. For quantitative predictive uncertainty assessment, Monte Carlo (MC)-Dropout is integrated with conformal prediction. This provides statistically valid prediction sets with entropy-based uncertainty visualization. The framework is evaluated on four medical imaging benchmark datasets: chest X-rays of COVID-19, Tuberculosis, Pneumonia, and retinal Optical Coherence Tomography (OCT) images. The proposed framework achieved strong classification performance with an AUC of 99.75% for COVID-19, 100% for Tuberculosis, 99.3% for Pneumonia chest X-rays, and 98.69% for retinal OCT images. Uncertainty-aware inference yields calibrated prediction sets with interpretable examples of uncertainty, showing transparency. The results demonstrate that bidirectional cross-attention with uncertainty quantification can improve performance and transparency in medical image classification.

医疗影像可解释性不确定性量化注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。