医学视觉语言模型的注意力图看似合理,实则常与真实诊断依据无关。
Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs

- 通过遮挡实验和放射科医生标注对比,检验注意力图是否真正指向关键病变区域。
- 所有测试模型的注意力图均未同时满足图像使用与注意力集中于关键区域的要求。
- 临床应用需依赖因果扰动而非视觉直观,现有热力图不可信。
注意力图和显著性热图被广泛用于解释医学视觉语言模型(VLM)在胸部X光片上的输出,但其是否真实反映驱动预测的图像证据尚未经过因果验证。本研究通过三种方式评估忠实性:在PadChest数据集(n=637)上计算与放射科医生标注框的重叠率,在CheXlocalize数据集(n=643)上测量归因质量与医生掩码的重合度,并利用16×16像素遮挡图记录哪些区域被隐藏时会改变预测结果。我们测试了三个MedGemma-4B变体、跨家族探针在LLaVA-RAD和Qwen3-VL-8B-Instruct上的表现,以及专用模型CheXagent-2-3b,以两个经胸片训练的分类器(DenseNet121、ResNet50)作为正向对照。只有当模型实际使用图像且注意力集中在遮挡后影响预测的区域时,热图才被认为是可信的。结果显示,无一被测VLM同时满足这两项标准。MedGemma和Qwen3-VL虽使用图像,但注意力与遮挡重要性呈负相关(rho < 0,95%置信区间全低于零);LLaVA-RAD注意力正相关,但模型近乎纯文本(99.1%文本一致性,因果质量接近零),相关性仅连接两个几乎为零的信号。注意力还未能捕捉到标注解剖结构:与真实区域的重叠始终低于随机或移位控制组,且无方法将超过22%的归因质量置于放射科医生掩码内。两个胸片分类器均通过所有指标,表明失败是特定于VLM热图,而非评估方法本身。这些热图虽具视觉说服力,却缺乏因果忠实性;临床解释必须结合受控定位指标与因果扰动,而非仅依赖视觉观察。
原文摘要 · Abstract (English)
Attention and saliency heatmaps are widely used to explain medical Vision-Language Model (VLM) outputs on chest X-rays, yet whether they truly highlight the image evidence driving predictions has not been causally tested. We audit faithfulness via overlap with radiologist bounding boxes on PadChest (n=637), attribution mass within radiologist masks on CheXlocalize (n=643), and 16x16 patch-occlusion maps that record which regions, when hidden, change the answer. We study three MedGemma-4B variants, cross-family probes on LLaVA-RAD and Qwen3-VL-8B-Instruct, and the specialist CheXagent-2-3b, with two CXR-trained classifiers (DenseNet121, ResNet50) as positive controls. A heatmap is faithful only if the model uses the image and attention concentrates on regions whose occlusion alters the prediction. No evaluated VLM meets both criteria. MedGemma and Qwen3-VL use the image, but attention anti-correlates with patch-occlusion importance (rho < 0 with 95% bootstrap CIs below zero). LLaVA-RAD's attention correlates positively, but the model is almost text-only (99.1% text-only agreement, near-zero causal mass), so correlation ties two near-zero signals. Attention also misses annotated anatomy: overlap with true regions never beats shifted or random controls, and no method places more than 22% of its mass inside radiologist masks. The two CXR classifiers pass all metrics, indicating the failure is specific to VLM heatmaps, not the evaluation. These heatmaps are visually reassuring but not faithful; clinical explanations require controlled localization metrics and causal perturbation, not visual inspection alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。