arXiv:2608.28207cs.CVcs.LG2026-08

用视觉大模型提升糖尿病视网膜病变诊断可解释性,效果优于传统方法。

Explainable Diabetic Retinopathy Classification Using Vision Foundation Models

论文配图:Explainable Diabetic Retinopathy Classification Using Vision Foundation Models
图 1 · 摘自论文原文
  • 采用DINOv2等大模型+LoRA微调,兼顾性能与参数效率。
  • 外部验证在APTOS数据集上达0.920的AUROC,内部达0.758。
  • 通过热力图与专家标注对比,证明模型关注点符合临床病灶位置。

糖尿病视网膜病变(DR)是导致可预防失明的主要原因,亟需准确且可信的自动化筛查。本研究探索了基于视觉基础模型的可解释性DR分类框架,采用多种迁移学习策略。评估了DINOv2、CLIP和Vision Transformer(ViT)三种主干网络,分别使用全微调、线性探针和低秩适应(LoRA)。模型在ODIR数据集上训练并内部评估,外部在APTOS上测试泛化能力。DINOv2-LoRA内部AUROC最高达0.758,而DINOv2全微调和ViT全微调在外部达到0.920的最高AUROC。经等距回归校准后,可靠性分析进一步验证模型输出稳定性。可解释性方面,使用Grad-CAM和HiResCAM生成注意力图,与IDRiD数据集上的专家标注病灶掩码进行对比,采用Dice、交并比(IoU)和指向游戏(Pointing Game)指标评估。结果表明,基础模型(尤其是DINOv2)能提供强大预测性能,同时LoRA为全微调提供了高效的替代方案;定量评估也证实模型关注区域与临床相关视网膜病变高度一致。

原文摘要 · Abstract (English)

Diabetic retinopathy (DR) is a major cause of preventable blindness, creating a need for accurate and trustworthy automated screening. This study investigates an explainable DR classification framework using vision foundation models and multiple transfer learning strategies. Three backbones, DINOv2, CLIP, and Vision Transformer (ViT), were evaluated using full fine-tuning, linear probing, and Low-Rank Adaptation (LoRA). Models were trained and internally evaluated on the ODIR dataset and externally evaluated on APTOS to assess generalization. DINOv2-LoRA achieved the highest internal AUROC of 0.758, while DINOv2 full fine-tuning and ViT full fine-tuning achieved the highest external AUROC of 0.920. Calibration was further assessed using reliability analysis after isotonic regression. For explainability, Grad-CAM and HiResCAM were evaluated against expert-annotated lesion masks from the IDRiD dataset using Dice, Intersection over Union (IoU), and Pointing Game metrics. The results demonstrate that foundation models, particularly DINOv2, can provide strong predictive performance, while LoRA offers a parameter-efficient alternative to full fine-tuning. Quantitative evaluation of explanation maps further supports the assessment of whether model attention corresponds to clinically relevant retinal lesions.

糖尿病视网膜病变视觉大模型可解释性迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。