分析了医学影像中Grad-CAM解释的可信度,发现其在Transformer模型上可靠性显著下降。
Seeing Isn't Always Believing: Analysis of Grad-CAM Faithfulness and Localization Reliability in Lung Cancer CT Classification
- 构建多维度评估框架,结合定位精度与扰动实验检验Grad-CAM可信度
- 卷积网络中Grad-CAM能准确定位肿瘤区域,但Vision Transformer效果大幅下降
- 提醒医疗AI领域谨慎使用可视化工具,需发展更可靠的可解释方法
可解释人工智能(XAI)技术如梯度加权类激活映射(Grad-CAM)已成为医学图像分析中揭示深度神经网络决策过程的重要工具。尽管广泛应用,其热图解释的忠实性与可靠性仍存争议。本研究以公开的IQ-OTH/NCCD数据集为基础,评估了五种典型架构:ResNet-50、ResNet-101、DenseNet-161、EfficientNet-B0和ViT-Base-Patch16-224,探究模型差异对Grad-CAM可解释性的影响。通过融合定位准确性、扰动忠实性与解释一致性构建定量评估框架。实验表明,大多数卷积网络中Grad-CAM能有效突出病灶区域,但视觉变换器(Vision Transformer)因非局部注意力机制导致解释忠实性显著下降。跨模型对比显示显著的显著性定位差异,表明Grad-CAM解释未必反映网络真实诊断依据。研究揭示了当前基于显著性的XAI方法在医学影像中的关键局限,强调需发展兼具计算合理性与临床意义的模型感知型可解释方法。研究呼吁在医疗AI中更审慎、严谨地应用可视化工具,重新思考对模型解释‘信任’的含义。
原文摘要 · Abstract (English)
Explainable Artificial Intelligence (XAI) techniques, such as Gradient-weighted Class Activation Mapping (Grad-CAM), have become indispensable for visualizing the reasoning process of deep neural networks in medical image analysis. Despite their popularity, the faithfulness and reliability of these heatmap-based explanations remain under scrutiny. This study critically investigates whether Grad-CAM truly represents the internal decision-making of deep models trained for lung cancer image classification. Using the publicly available IQ-OTH/NCCD dataset, we evaluate five representative architectures: ResNet-50, ResNet-101, DenseNet-161, EfficientNet-B0, and ViT-Base-Patch16-224, to explore model-dependent variations in Grad-CAM interpretability. We introduce a quantitative evaluation framework that combines localization accuracy, perturbation-based faithfulness, and explanation consistency to assess Grad-CAM reliability across architectures. Experimental findings reveal that while Grad-CAM effectively highlights salient tumor regions in most convolutional networks, its interpretive fidelity significantly degrades for Vision Transformer models due to non-local attention behavior. Furthermore, cross-model comparisons indicate substantial variability in saliency localization, implying that Grad-CAM explanations may not always correspond to the true diagnostic evidence used by the networks. This work exposes critical limitations of current saliency-based XAI approaches in medical imaging and emphasizes the need for model-aware interpretability methods that are both computationally sound and clinically meaningful. Our findings aim to inspire a more cautious and rigorous adoption of visual explanation tools in medical AI, urging the community to rethink what it truly means to "trust" a model's explanation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。