用多模态模型分析眼底图像四象限,实现糖尿病视网膜病变的可解释诊断。
Quadrant Segmentation VLM with Few-Shot Adaptation and OCT Learning-based Explainability Methods for Diabetic Retinopathy
- 基于视觉语言模型,通过四象限分析眼底与OCT图像进行少样本自适应。
- 在3000张眼底图和1000张OCT图上实现高精度分类与可解释性输出。
- 生成热力图展示关键区域贡献,适合临床筛查与医生决策支持。
糖尿病视网膜病变(DR)是全球失明的主要原因,需早期发现以保护视力。由于医生资源有限,常导致漏诊。现有AI模型依赖病灶分割提升可解释性,但人工标注病灶对临床不现实。医生更关注模型判断依据,而非仅定位病灶。当前模型多为单模态,解释效果有限。本文提出一种新型多模态可解释模型,结合视觉语言模型(VLM)与少样本学习,模仿眼科医生思路,分析眼底图像中病灶在四个象限的分布,实现自然语言描述的定量检测。模型生成配对的Grad-CAM热力图,可视化神经元在眼底与OCT图像上的权重分布,明确显示影响严重程度判断的关键区域。基于3,000张眼底图像和1,000张OCT图像的数据集,该方法克服了现有诊断中的关键局限,为筛查、治疗和研究提供实用且全面的工具。
原文摘要 · Abstract (English)
Diabetic Retinopathy (DR) is a leading cause of vision loss worldwide, requiring early detection to preserve sight. Limited access to physicians often leaves DR undiagnosed. To address this, AI models utilize lesion segmentation for interpretability; however, manually annotating lesions is impractical for clinicians. Physicians require a model that explains the reasoning for classifications rather than just highlighting lesion locations. Furthermore, current models are one-dimensional, relying on a single imaging modality for explainability and achieving limited effectiveness. In contrast, a quantitative-detection system that identifies individual DR lesions in natural language would overcome these limitations, enabling diverse applications in screening, treatment, and research settings. To address this issue, this paper presents a novel multimodal explainability model utilizing a VLM with few-shot learning, which mimics an ophthalmologist's reasoning by analyzing lesion distributions within retinal quadrants for fundus images. The model generates paired Grad-CAM heatmaps, showcasing individual neuron weights across both OCT and fundus images, which visually highlight the regions contributing to DR severity classification. Using a dataset of 3,000 fundus images and 1,000 OCT images, this innovative methodology addresses key limitations in current DR diagnostics, offering a practical and comprehensive tool for improving patient outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。