通过疾病感知提示增强医学图像定位准确率
Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding
- 用视觉模型解释图生成提示,聚焦病灶区域
- 在三组胸片数据集上提升20.74%定位准确率
- 无需像素级标注,适合临床可解释性需求
视觉定位(VG)是识别图像中与特定文本描述相关区域的能力。在医学影像中,VG通过突出对应文本描述的病理特征,提升模型可解释性,增强临床对深度学习模型的信任。现有模型因注意力机制低效及细粒度标记表示不足,难以将文本与病灶区域关联。本文实证发现:第一,当前视觉语言模型对背景标记赋予高范数,分散模型对病灶的关注;第二,用于跨模态学习的全局标记不能代表局部病灶标记,阻碍文本与病灶标记间的关联。为此,我们提出简单而有效的疾病感知提示(DAP)方法,利用视觉语言模型的可解释性图识别合适图像特征,放大病灶相关区域并抑制背景干扰。无需额外像素级标注,DAP在三个主要胸片数据集上相比最先进方法提升20.74%的视觉定位准确率。
原文摘要 · Abstract (English)
Visual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability by highlighting relevant pathological features corresponding to textual descriptions, improving model transparency and trustworthiness for wider adoption of deep learning models in clinical practice. Current models struggle to associate textual descriptions with disease regions due to inefficient attention mechanisms and a lack of fine-grained token representations. In this paper, we empirically demonstrate two key observations. First, current VLMs assign high norms to background tokens, diverting the model's attention from regions of disease. Second, the global tokens used for cross-modal learning are not representative of local disease tokens. This hampers identifying correlations between the text and disease tokens. To address this, we introduce simple, yet effective Disease-Aware Prompting (DAP) process, which uses the explainability map of a VLM to identify the appropriate image features. This simple strategy amplifies disease-relevant regions while suppressing background interference. Without any additional pixel-level annotations, DAP improves visual grounding accuracy by 20.74% compared to state-of-the-art methods across three major chest X-ray datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。