用知识图谱增强医学图文问答,提升诊断关联与自由回答能力
KG-CMI: Knowledge graph enhanced cross-Mamba interaction for medical visual question answering
- 通过知识图谱融合医学专业知识,建立病灶特征与疾病信息的关联
- 在VQA-RAD、SLAKE、OVQA三个数据集上超越现有最佳方法
- 支持开放式答案生成,适合临床决策与远程医疗场景
医学视觉问答(Med-VQA)是临床决策支持和远程医疗中的关键多模态任务。现有方法未能充分挖掘领域特定医学知识,难以准确将医学图像中的病灶特征与关键诊断标准关联。此外,基于分类的方法通常依赖预定义答案集,将Med-VQA视为简单分类问题会限制其对自由形式答案的适应性,并可能忽略答案中的详细语义信息。为此,我们提出知识图谱增强的跨Mamba交互框架(KG-CMI),包含细粒度跨模态特征对齐(FCFA)、知识图嵌入(KGE)、跨模态交互表示(CMIR)和自由形式答案增强多任务学习(FAMT)模块。KG-CMI通过图结构有效整合专业医学知识,学习图像与文本间的跨模态表示,建立病灶特征与疾病知识之间的联系。同时,FAMT利用开放问题中的辅助知识,提升模型对开放式Med-VQA的能力。实验结果表明,KG-CMI在三个Med-VQA数据集(VQA-RAD、SLAKE、OVQA)上均优于现有最先进方法。此外,我们还进行了可解释性实验,进一步验证了该框架的有效性。
原文摘要 · Abstract (English)
Medical visual question answering (Med-VQA) is a crucial multimodal task in clinical decision support and telemedicine. Recent methods fail to fully leverage domain-specific medical knowledge, making it difficult to accurately associate lesion features in medical images with key diagnostic criteria. Additionally, classification-based approaches typically rely on predefined answer sets. Treating Med-VQA as a simple classification problem limits its ability to adapt to the diversity of free-form answers and may overlook detailed semantic information in those answers. To address these challenges, we propose a knowledge graph enhanced cross-Mamba interaction (KG-CMI) framework, which consists of a fine-grained cross-modal feature alignment (FCFA) module, a knowledge graph embedding (KGE) module, a cross-modal interaction representation (CMIR) module, and a free-form answer enhanced multi-task learning (FAMT) module. The KG-CMI learns cross-modal feature representations for images and texts by effectively integrating professional medical knowledge through a graph, establishing associations between lesion features and disease knowledge. Moreover, FAMT leverages auxiliary knowledge from open-ended questions, improving the model's capability for open-ended Med-VQA. Experimental results demonstrate that KG-CMI outperforms existing state-of-the-art methods on three Med-VQA datasets, i.e., VQA-RAD, SLAKE, and OVQA. Additionally, we conduct interpretability experiments to further validate the framework's effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。