用医学知识分解提升视觉模型对病灶的定位能力
Enhancing Abnormality Grounding for Vision Language Models with Knowledge Descriptions
- 将医学术语拆解为基础属性和视觉模式,增强文本与图像对齐
- 仅用1.5%数据训练,性能媲美70亿参数大模型
- 对已知和未知病灶均有良好泛化效果,适合医疗视觉任务
视觉语言模型(VLMs)在视觉定位任务中表现优异,但在医学领域,尤其是医学图像中的异常检测与定位方面仍研究不足。主要挑战在于医学术语复杂抽象,难以与视觉特征直接关联。本文提出一种新方法,通过分解医学知识来提升VLM在医学异常检测与定位中的表现。不直接引导模型识别特定异常,而是将医学概念拆解为基础属性和常见视觉模式,从而加强文本描述与视觉特征之间的对齐,提升异常识别与定位能力。我们在0.23B规模的Florence-2基础模型上验证该方法,结果表明,尽管仅使用同类模型训练数据的1.5%,其性能可媲美参数量达7B的LLaVA-based医学VLM。实验还证明该方法在已知及未见过的异常上均有效,展现出强泛化能力。
原文摘要 · Abstract (English)
Visual Language Models (VLMs) have demonstrated impressive capabilities in visual grounding tasks. However, their effectiveness in the medical domain, particularly for abnormality detection and localization within medical images, remains underexplored. A major challenge is the complex and abstract nature of medical terminology, which makes it difficult to directly associate pathological anomaly terms with their corresponding visual features. In this work, we introduce a novel approach to enhance VLM performance in medical abnormality detection and localization by leveraging decomposed medical knowledge. Instead of directly prompting models to recognize specific abnormalities, we focus on breaking down medical concepts into fundamental attributes and common visual patterns. This strategy promotes a stronger alignment between textual descriptions and visual features, improving both the recognition and localization of abnormalities in medical images.We evaluate our method on the 0.23B Florence-2 base model and demonstrate that it achieves comparable performance in abnormality grounding to significantly larger 7B LLaVA-based medical VLMs, despite being trained on only 1.5% of the data used for such models. Experimental results also demonstrate the effectiveness of our approach in both known and previously unseen abnormalities, suggesting its strong generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。