用显著性引导注意力,让模型跨时间看懂胸部X光变化。
Saliency Guided Longitudinal Medical Visual Question Answering

- 通过关键词生成病变关注区域,强制模型跨时间保持一致注意力。
- 在Medical-Diff-VQA上达到主流指标表现,无需放射科预训练。
- 适合需要可解释性医疗视觉问答的临床研究与开发人员。
纵向医学视觉问答(Diff-VQA)需对比不同时期的影像对并回答关于临床变化的问题。在此任务中,差异信号和跨时间视觉焦点的一致性比单图绝对发现更具信息量。我们提出一种基于显著性的编码器-解码器框架,将后验显著性转化为可操作的监督信号。模型首先执行轻量级近恒等仿射预对齐,以减少就诊间的冗余运动;随后进行本周期内两步循环:第一步从答案中提取医学相关关键词,并在两张图像上生成关键词条件的Grad-CAM,获取病灶聚焦的显著性;第二步将共享显著性掩码应用于两个时间点,生成最终答案。该机制闭环连接语言与视觉,使关键术语也指导模型关注位置,强制对应解剖结构的空间一致性注意力。在Medical-Diff-VQA数据集上,该方法在BLEU、ROUGE-L、CIDEr和METEOR指标上表现优异,且具备内在可解释性。值得注意的是,主干网络和解码器均使用通用领域预训练,未采用放射科特定预训练,凸显其实用性与迁移能力。结果支持在轻微预对齐基础上,采用显著性条件生成是一种合理的纵向推理框架。
原文摘要 · Abstract (English)
Longitudinal medical visual question answering (Diff-VQA) requires comparing paired studies from different time points and answering questions about clinically meaningful changes. In this setting, the difference signal and the consistency of visual focus across time are more informative than absolute single-image findings. We propose a saliency-guided encoder-decoder for chest X-ray Diff-VQA that turns post-hoc saliency into actionable supervision. The model first performs a lightweight near-identity affine pre-alignment to reduce nuisance motion between visits. It then executes a within-epoch two-step loop: step 1 extracts a medically relevant keyword from the answer and generates keyword-conditioned Grad-CAM on both images to obtain disease-focused saliency; step 2 applies the shared saliency mask to both time points and generates the final answer. This closes the language-vision loop so that the terms that matter also guide where the model looks, enforcing spatially consistent attention on corresponding anatomy. On Medical-Diff-VQA, the approach attains competitive performance on BLEU, ROUGE-L, CIDEr, and METEOR while providing intrinsic interpretability. Notably, the backbone and decoder are general-domain pretrained without radiology-specific pretraining, highlighting practicality and transferability. These results support saliency-conditioned generation with mild pre-alignment as a principled framework for longitudinal reasoning in medical VQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。