通过标注视觉感兴趣区域,提升医学图像问答模型理解能力
R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest
- 将医学标注的视觉区域作为先验知识注入图像空间
- 在4个标准数据集上优于当前最优方法
- 适合医疗AI研究者和医学图像分析开发者
人工智能在医学视觉问答(Med-VQA)领域取得显著进展,但现有研究多采用整体图像理解,忽略可能包含关键信息的视觉区域。本文提出R-LLaVA,通过CLIP将简单的医学标注(如边界框)作为先验知识直接融入图像空间,并在训练中输入到LLaVA模型,以增强对生物医学问题的理解。在四个标准Med-VQA数据集上的实验表明,R-LLaVA优于现有最先进方法。此外,为验证模型视觉理解能力,本文构建了一个新的多选医学视觉理解数据集,证实聚焦视觉感兴趣区域能有效提升生物医学VQA性能。
原文摘要 · Abstract (English)
Artificial intelligence has made significant strides in medical visual question answering (Med-VQA), yet prevalent studies often interpret images holistically, overlooking the visual regions of interest that may contain crucial information, potentially aligning with a doctor's prior knowledge that can be incorporated with minimal annotations (e.g., bounding boxes). To address this gap, this paper introduces R-LLaVA, designed to enhance biomedical VQA understanding by integrating simple medical annotations as prior knowledge directly into the image space through CLIP. These annotated visual regions of interest are then fed into the LLaVA model during training, aiming to enrich the model's understanding of biomedical queries. Experimental evaluation on four standard Med-VQA datasets demonstrates R-LLaVA's superiority over existing state-of-the-art (SoTA) methods. Additionally, to verify the model's capability in visual comprehension, a novel multiple-choice medical visual understanding dataset is introduced, confirming the positive impact of focusing on visual regions of interest in advancing biomedical VQA understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。