深度学习模型诊断肺癌准确率高,但推理逻辑差异大,需独立评估解释性。
Trusting Right Predictions for Wrong Reasons: A LIME Based Analysis of Deep Learning Interpretability in Lung Cancer Diagnosis

- 用LIME分析三种模型的决策依据,发现预测一致但解释差异显著。
- 三模型准确率超93%,但解释相关性低于0.26,说明推理不一致。
- 错误判断常因关注肺外区域,正确判断聚焦肺组织,适合临床可信度研究。
肺癌是导致癌症死亡的主要原因,每年约有250万新发病例和180万死亡病例,可靠诊断具有重要临床意义。尽管深度学习在肺癌分类中表现优异,但评估仍以预测准确率为重心,对其决策过程缺乏深入分析。本研究对比了三种架构不同的模型:卷积神经网络(CNN)、预训练的ResNet50和视觉变换器(ViT),均在IQ-OTH/NCCD肺部CT数据集上训练。采用局部可解释模型无关解释(LIME)分析模型推理过程,并引入双相关性框架衡量模型对预测与解释的一致性。所有模型表现优异:ResNet50准确率达98.61%,CNN为97.91%,ViT为93.75%,且所有模型的ROC-AUC均为0.99。模型间预测相关性超过0.99,表明输出高度一致;但解释相关性均低于0.26,显示其关注图像区域差异显著。进一步分析误判样本发现,错误预测多关联肺外区域注意力,而正确预测则集中于肺实质内部。结果表明,预测一致性不能代表推理一致性,解释性评估必须作为临床人工智能系统中独立于预测性能的验证标准。
原文摘要 · Abstract (English)
Lung cancer is the leading cause of cancer-related mortality, with approximately 2.5 million new cases and 1.8 million deaths annually, making reliable diagnosis a clinical priority. Although deep learning models have achieved strong performance in lung cancer classification, evaluation has largely focused on predictive accuracy, leaving their decision-making processes insufficiently examined. This study compares three architecturally distinct models: a Convolutional Neural Network (CNN), a pretrained ResNet50, and a Vision Transformer (ViT), trained on the IQ-OTH/NCCD lung cancer CT dataset. Local Interpretable Model-Agnostic Explanations (LIME) were applied to investigate model reasoning. In addition to standard performance metrics, a dual-correlation framework was introduced to measure both prediction agreement and explanation agreement across model pairs. All three models achieved strong classification performance, with ResNet50 attaining 98.61% accuracy, CNN 97.91%, and ViT 93.75%, while all achieved ROC-AUC scores of 0.99. Prediction correlations exceeded 0.99 across all model pairs, indicating highly consistent outputs. However, LIME explanation correlations remained below 0.26, revealing substantial differences in the image regions used to reach those predictions. Analysis of misclassified samples further identified a consistent spatial pattern: incorrect predictions were associated with attention outside the lung parenchyma, whereas correct predictions focused primarily within lung regions. These findings demonstrate that prediction agreement is a poor proxy for reasoning consistency, and that interpretability evaluation must be treated as an independent validation criterion alongside predictive performance in clinical AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。