arXiv:2508.15168cs.CV2025-08被引 4

用可解释的视觉语言模型,让糖尿病视网膜病变诊断更准更透明。

XDR-LVLM: An Explainable Vision-Language Large Model for Diabetic Retinopathy Diagnosis

  • 融合医学视觉编码器与多任务提示工程,实现病灶特征深度理解。
  • 在DDR数据集上诊断准确率达84.55%,关键病灶识别F1达66.88%。
  • 生成自然语言报告,适合医生辅助诊断与AI可解释性研究者使用。

糖尿病视网膜病变(DR)是全球致盲主因,需早期精准诊断。尽管深度学习模型在检测中表现良好,但其黑箱特性限制了临床应用。为此,我们提出XDR-LVLM(基于视觉语言大模型的可解释糖尿病视网膜病变诊断),利用视觉语言大模型(LVLM)实现高精度诊断并生成自然语言解释。该框架包含专用医学视觉编码器、LVLM核心,并采用多任务提示工程与多阶段微调,深入理解眼底图像中的病理特征,生成包含疾病严重程度分级、关键病灶识别(如出血、渗出、微动脉瘤)及对应诊断依据的综合报告。在DDR数据集上的实验表明,XDR-LVLM在疾病诊断上达到84.55%平衡准确率和79.92%的F1分数,在病灶检测中取得77.95%平衡准确率与66.88%的F1分数。人工评估验证了生成解释的流畅性、准确性与临床实用性,证明其能有效弥合自动化诊断与临床需求之间的差距。

原文摘要 · Abstract (English)

Diabetic Retinopathy (DR) is a major cause of global blindness, necessitating early and accurate diagnosis. While deep learning models have shown promise in DR detection, their black-box nature often hinders clinical adoption due to a lack of transparency and interpretability. To address this, we propose XDR-LVLM (eXplainable Diabetic Retinopathy Diagnosis with LVLM), a novel framework that leverages Vision-Language Large Models (LVLMs) for high-precision DR diagnosis coupled with natural language-based explanations. XDR-LVLM integrates a specialized Medical Vision Encoder, an LVLM Core, and employs Multi-task Prompt Engineering and Multi-stage Fine-tuning to deeply understand pathological features within fundus images and generate comprehensive diagnostic reports. These reports explicitly include DR severity grading, identification of key pathological concepts (e.g., hemorrhages, exudates, microaneurysms), and detailed explanations linking observed features to the diagnosis. Extensive experiments on the Diabetic Retinopathy (DDR) dataset demonstrate that XDR-LVLM achieves state-of-the-art performance, with a Balanced Accuracy of 84.55% and an F1 Score of 79.92% for disease diagnosis, and superior results for concept detection (77.95% BACC, 66.88% F1). Furthermore, human evaluations confirm the high fluency, accuracy, and clinical utility of the generated explanations, showcasing XDR-LVLM's ability to bridge the gap between automated diagnosis and clinical needs by providing robust and interpretable insights.

糖尿病视网膜病变可解释AI视觉语言模型医疗诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。