arXiv:2510.12021cs.CV2025-10中稿 · Workshop on Interp…被引 6

对比多种视觉Transformer在医学影像中的可解释性,发现DINO+Grad-CAM效果最佳。

Evaluating the Explainability of Vision Transformers in Medical Imaging

  • 用Grad-CAM和注意力回滚评估ViT类模型的解释能力
  • DINO+Grad-CAM生成最精准、最忠实的热力图,定位准确
  • 即使分类错误也能指出关键病理特征,适合临床可信应用

理解模型决策对医学影像至关重要,直接影响临床信任与应用。视觉Transformer(ViTs)在诊断影像中表现优异,但其复杂的注意力机制带来解释难题。本研究评估了ViT、DeiT、DINO和Swin Transformer四种架构及预训练策略,采用梯度注意力回滚(Gradient Attention Rollout)和Grad-CAM方法,在外周血细胞分类和乳腺超声图像分类两个任务上进行定量与定性分析。结果表明,DINO结合Grad-CAM在不同数据集上均提供最忠实且空间定位精确的解释。Grad-CAM持续生成具有类别判别力和空间精度的热力图,而注意力回滚则产生更分散的激活。即使在误分类情况下,该组合仍能突出提示模型误判的临床形态学特征。研究提升了模型透明度,支持将ViTs可靠、可解释地集成至关键医疗诊断流程。

原文摘要 · Abstract (English)

Understanding model decisions is crucial in medical imaging, where interpretability directly impacts clinical trust and adoption. Vision Transformers (ViTs) have demonstrated state-of-the-art performance in diagnostic imaging; however, their complex attention mechanisms pose challenges to explainability. This study evaluates the explainability of different Vision Transformer architectures and pre-training strategies - ViT, DeiT, DINO, and Swin Transformer - using Gradient Attention Rollout and Grad-CAM. We conduct both quantitative and qualitative analyses on two medical imaging tasks: peripheral blood cell classification and breast ultrasound image classification. Our findings indicate that DINO combined with Grad-CAM offers the most faithful and localized explanations across datasets. Grad-CAM consistently produces class-discriminative and spatially precise heatmaps, while Gradient Attention Rollout yields more scattered activations. Even in misclassification cases, DINO with Grad-CAM highlights clinically relevant morphological features that appear to have misled the model. By improving model transparency, this research supports the reliable and explainable integration of ViTs into critical medical diagnostic workflows.

视觉Transformer医学影像可解释性Grad-CAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。