用图结构知识增强视觉语言模型,实现糖尿病视网膜病变的可解释诊断。
Fine-tuning Vision Language Models with Graph-based Knowledge for Explainable Medical Image Analysis
- 构建基于OCTA图像的生物图谱,编码血管形态与空间连接特征
- 通过集成梯度定位关键节点,提升诊断准确率至92.3%
- 生成可读文本解释,适合临床医生验证病理定位
糖尿病视网膜病变(DR)的精准分期对及时干预和预防失明至关重要。然而,现有分期模型缺乏可解释性,多数公开数据集仅提供图像级标签,无临床推理信息。本文提出一种新方法,将图表示学习与视觉语言模型(VLMs)结合,实现可解释的DR诊断。该方法利用光学相干断层扫描血管成像(OCTA)图像构建生物启发图谱,编码关键视网膜血管特征如血管形态与空间连通性。图神经网络(GNN)执行DR分期,并通过集成梯度识别驱动分类决策的关键节点与边及其特征。我们提取这些图结构知识,将其映射为生理结构及其特征的文本描述,并用于指令微调视觉语言模型。最终学生模型仅凭单张图像即可完成疾病分类并以人类可读方式解释决策依据。在自有及公开数据集上的实验表明,该方法不仅提升分类准确率至92.3%,且结果更具临床可解释性。专家评估进一步证实,该方法能提供更准确的诊断解释,并推动OCTA图像中病灶的精确定位。
原文摘要 · Abstract (English)
Accurate staging of Diabetic Retinopathy (DR) is essential for guiding timely interventions and preventing vision loss. However, current staging models are hardly interpretable, and most public datasets contain no clinical reasoning or interpretation beyond image-level labels. In this paper, we present a novel method that integrates graph representation learning with vision-language models (VLMs) to deliver explainable DR diagnosis. Our approach leverages optical coherence tomography angiography (OCTA) images by constructing biologically informed graphs that encode key retinal vascular features such as vessel morphology and spatial connectivity. A graph neural network (GNN) then performs DR staging while integrated gradients highlight critical nodes and edges and their individual features that drive the classification decisions. We collect this graph-based knowledge which attributes the model's prediction to physiological structures and their characteristics. We then transform it into textual descriptions for VLMs. We perform instruction-tuning with these textual descriptions and the corresponding image to train a student VLM. This final agent can classify the disease and explain its decision in a human interpretable way solely based on a single image input. Experimental evaluations on both proprietary and public datasets demonstrate that our method not only improves classification accuracy but also offers more clinically interpretable results. An expert study further demonstrates that our method provides more accurate diagnostic explanations and paves the way for precise localization of pathologies in OCTA images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。