轻量级医学视觉语言模型,精准聚焦病灶区域提升诊断准确率
RoiMAM: Region-of-Interest Medical Attention Model for Efficient Vision-Language Understanding

- 通过无训练的病灶区域生成与语义抑制,聚焦关键病变区域
- 模型规模不足20%,在SLAKE上提升2%准确率,PMC-VQA提升4.6%
- 无需额外训练参数,适合临床部署的高效医疗问答系统
视觉语言模型(VLM)通过联合解析图像与文本,推动医学视觉问答(MedVQA)发展。然而,现有模型通常依赖大型架构和封闭集答案,限制了效率与临床适用性。为此,我们提出RoiMAM,一种高效VLM。它融合无训练的病灶区域生成模块与语义选择性抑制机制,聚焦病变相关区域,并引入文本提示增强模块,提供模态特异性上下文而无需增加训练参数。相比广泛使用的MedVInT-TD模型,RoiMAM在模型规模小于20%的情况下,于SLAKE数据集上准确率提升约2%,在PMC-VQA上提升4.6%。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) facilitate medical visual question answering (MedVQA) by jointly interpreting images and text. However, existing models typically depend on large architectures and closed-set answers, which limits their efficiency and potential clinical applicability. To overcome these shortcomings, we introduce RoiMAM, an efficient VLM. It integrates a training-free ROI Generation Module with Semantic Selective Suppression to focus on lesion-relevant regions, alongside a Text Prompt Enhancer module that provides modality-specific context without introducing training parameters. Compared to the widely used MedVInT-TD model, our design achieves efficient and accurate diagnosis at less than 20\% of the model size, while improving accuracy by approximately 2% on SLAKE and 4.6% on PMC-VQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。