arXiv:2603.13403cs.CV2026-03被引 1

用排序感知提示提升眼底图像糖尿病视网膜病变分级准确率

Diabetic Retinopathy Grading with CLIP-based Ranking-Aware Adaptation:A Comparative Study on Fundus Image

  • 设计排序感知提示,捕捉病变严重程度的有序性
  • 最高准确率达93.42%,重症病例召回率显著提升
  • 适合临床筛查场景,对重症检测尤为可靠

糖尿病视网膜病变(DR)是可预防失明的主要原因,自动化眼底图像分级在大规模筛查中具有重要意义。本文研究三种基于CLIP的方法用于五级DR严重程度分级:(1) 使用提示工程的零样本基线,(2) 增强CBAM注意力的混合全卷积网络-CLIP模型,(3) 编码DR进展顺序结构的排序感知提示模型。在合并APTOS 2019与Messidor-2数据集(n=5,406)上进行训练与评估,通过重采样和类别特异性最优阈值处理类别不平衡问题。实验表明,排序感知模型达到最高总体准确率(93.42%,AUROC 0.9845),且对临床关键的重症病例具有强召回能力;混合FCN-CLIP模型(92.49%,AUROC 0.99)在检测增殖性DR方面表现优异。两者均显著优于零样本基线(55.17%,AUROC 0.75)。分析各方法互补优势,并讨论其在筛查场景中的实际意义。

原文摘要 · Abstract (English)

Diabetic retinopathy (DR) is a leading cause of preventable blindness, and automated fundus image grading can play an important role in large-scale screening. In this work, we investigate three CLIP-based approaches for five-class DR severity grading: (1) a zero-shot baseline using prompt engineering, (2) a hybrid FCN-CLIP model augmented with CBAM attention, and (3) a ranking-aware prompting model that encodes the ordinal structure of DR progression. We train and evaluate on a combined dataset of APTOS 2019 and Messidor-2 (n=5,406), addressing class imbalance through resampling and class-specific optimal thresholding. Our experiments show that the ranking-aware model achieves the highest overall accuracy (93.42%, AUROC 0.9845) and strong recall on clinically critical severe cases, while the hybrid FCN-CLIP model (92.49%, AUROC 0.99) excels at detecting proliferative DR. Both substantially outperform the zero-shot baseline (55.17%, AUROC 0.75). We analyze the complementary strengths of each approach and discuss their practical implications for screening contexts.

糖尿病视网膜病变眼底图像CLIP分级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。