用跨模态知识蒸馏提升医疗图像问答的对齐与理解能力
ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images
- 通过跨模态知识蒸馏强化图像与文本的对齐
- 在多个医疗问答数据集上达到当前最佳性能
- 适合缺乏医学先验知识的多任务医疗视觉模型
医学视觉问答(Med-VQA)是通用视觉问答领域中的关键且具有挑战性的子任务。尽管通用视觉问答取得显著进展,多模态大语言模型(MLLMs)在处理多任务场景时仍存在明显局限,主要表现为空间定位错误和医学图像误读,根源在于图像-文本对齐不足及医学领域知识缺失。为此,我们提出跨模态临床知识蒸馏框架(ClinKD),旨在增强图像-文本对齐并建立更有效的医学知识迁移机制,使MLLM在无医学先验知识的情况下也能表现更优。大量实验表明,ClinKD在多个具有挑战性的Med-VQA数据集上均达到先进水平。结果证实,该方法不仅显著提升图像-文本对齐效果,还能有效帮助MLLM适应医学知识。代码已开源:https://github.com/overloadedHenry/ClinKD。
原文摘要 · Abstract (English)
Medical Visual Question Answering (Med-VQA) represents a critical and challenging subtask within the general VQA domain. Despite significant advancements in general VQA, multimodal large language models (MLLMs) still exhibit substantial limitations when handling multi-task VQA scenarios. These limitations manifest through erroneous spatial localization and misinterpretation of medical images, which primarily arise from two fundamental issues: inadequate image-text alignment and insufficient domain-specified knowledge for medical applications. To address these issues, we introduce the Cross-Modal Clinical Knowledge Distiller (ClinKD), an innovative framework designed to enhance image-text alignment and establish more effective medical knowledge transformation mechanisms, which enables MLLMs to perform better even when lacking prior medical knowledge. Our extensive experimental evaluations demonstrate that the ClinKD achieves state-of-the-art performance on several datasets which are challenging for Med-VQA task. The results indicate that our approach not only significantly improves image-text alignment but also effectively enables MLLMs to adapt to the medical knowledge. The source code for ClinKD is available at: https://github.com/overloadedHenry/ClinKD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。