arXiv:2503.09535cs.CVcs.AI2025-03中稿 · publication in MIC…被引 18

对比注意力图与其它方法,发现其在医学影像解释中效果有限。

Evaluating Visual Explanations of Attention Maps for Transformer-based Medical Imaging

  • 用四个医学数据集测试注意力图的解释能力
  • 注意力图虽优于GradCAM,但不如专用可解释方法
  • 其解释效果依赖场景,难满足临床决策需求

尽管视觉变换器(ViTs)在医学影像任务中表现优异,但仍面临可解释性问题,与以往卷积神经网络类似。近期研究认为,作为决策过程一部分的注意力图可能通过识别影响预测的区域来缓解此问题,尤其在自监督预训练模型中。本文在四个医学影像数据集上进行大规模实验,涵盖结肠息肉、乳腺肿瘤、食管炎症及骨骨折与植入物的识别任务,使用多种监督和自监督预训练的ViT模型,比较注意力图与其他常用解释方法。结果表明,尽管注意力图在特定条件下有潜力,且通常优于GradCAM,但仍被专为变压器设计的可解释性方法超越。研究指出,注意力图的可解释性效能具有上下文依赖性,其提供的信息并不总是足够全面,难以支撑可靠的医学决策。

原文摘要 · Abstract (English)

Although Vision Transformers (ViTs) have recently demonstrated superior performance in medical imaging problems, they face explainability issues similar to previous architectures such as convolutional neural networks. Recent research efforts suggest that attention maps, which are part of decision-making process of ViTs can potentially address the explainability issue by identifying regions influencing predictions, especially in models pretrained with self-supervised learning. In this work, we compare the visual explanations of attention maps to other commonly used methods for medical imaging problems. To do so, we employ four distinct medical imaging datasets that involve the identification of (1) colonic polyps, (2) breast tumors, (3) esophageal inflammation, and (4) bone fractures and hardware implants. Through large-scale experiments on the aforementioned datasets using various supervised and self-supervised pretrained ViTs, we find that although attention maps show promise under certain conditions and generally surpass GradCAM in explainability, they are outperformed by transformer-specific interpretability methods. Our findings indicate that the efficacy of attention maps as a method of interpretability is context-dependent and may be limited as they do not consistently provide the comprehensive insights required for robust medical decision-making.

可解释AI视觉Transformer医学影像注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。