用医学知识增强空间提示,提升医疗图文定位精度。
Enhancing Medical Visual Grounding via Knowledge-guided Spatial Prompts
- 引入医学知识嵌入的提示策略,强化模型空间感知能力。
- 在四个基准上实现AP50提升3.0%,mIoU提升2.6%。
- 无需额外文本推理开销,适合临床辅助诊断场景。
医学视觉定位(MVG)旨在从自由文本放射科报告中识别具有诊断意义的短语,并在医学图像中定位其对应区域,为临床决策提供可解释的视觉证据。尽管近期视觉语言模型(VLMs)展现出出色的多模态推理能力,但其定位仍缺乏足够的空间精度,主要源于仅依赖隐式嵌入时缺乏显式定位先验。本文从注意力视角分析该限制,提出KnowMVG框架——一种基于知识先验与全局-局部注意力增强的MVG方法,在解码阶段显式加强空间意识。具体而言,我们设计了一种知识增强提示策略,将与短语相关的医学知识编码为紧凑嵌入,并引入全局-局部注意力机制,联合利用粗粒度全局信息与精细局部线索,以指导精准区域定位。该设计在不增加额外文本推理开销的前提下,弥合高层语义理解与细粒度视觉感知之间的鸿沟。在四个MVG基准上的大量实验表明,KnowMVG持续优于现有方法,在AP50上提升3.0%,在mIoU上提升2.6%。定性分析与消融研究进一步验证了各组件的有效性。
原文摘要 · Abstract (English)
Medical Visual Grounding (MVG) aims to identify diagnostically relevant phrases from free-text radiology reports and localize their corresponding regions in medical images, providing interpretable visual evidence to support clinical decision-making. Although recent Vision-Language Models (VLMs) exhibit promising multimodal reasoning ability, their grounding remains insufficient spatial precision, largely due to a lack of explicit localization priors when relying solely on latent embeddings. In this work, we analyze this limitation from an attention perspective and propose KnowMVG, a Knowledge-prior and global-local attention enhancement framework for MVG in VLMs that explicitly strengthens spatial awareness during decoding. Specifically, we present a knowledge-enhanced prompting strategy that encodes phrase related medical knowledge into compact embeddings, together with a global-local attention that jointly leverages coarse global information and refined local cues to guide precise region localization. localization. This design bridges high-level semantic understanding and fine-grained visual perception without introducing extra textual reasoning overhead. Extensive experiments on four MVG benchmarks demonstrate that our KnowMVG consistently outperforms existing approaches, achieving gains of 3.0% in AP50 and 2.6% in mIoU over prior state-of-the-art methods. Qualitative and ablation studies further validate the effectiveness of each component.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。