对比大模型与轻量CNN在糖尿病黄斑水肿检测中的表现
Evaluating Fundus-Specific Foundation Models for Diabetic Macular Edema Detection
- 用RETFound、FLAIR和EfficientNet-B0对比不同训练策略
- 轻量CNN在多数数据集上优于大模型,最高AUC领先0.05
- 大模型零样本表现不错,但微调后优势不明显适合小数据场景
糖尿病性黄斑水肿(DME)是糖尿病视网膜病变患者视力丧失的主要原因。尽管深度学习在从眼底图像自动检测该病方面表现出潜力,但受限于标注数据稀缺,实际应用仍具挑战。基础模型(FM)成为潜在解决方案。本文系统比较了用于视网膜图像的两种主流基础模型——RETFound和FLAIR,以及EfficientNet-B0骨干网络,在IDRiD、MESSIDOR-2和OEFI三个数据集上的不同训练与评估设置下的表现。结果显示,尽管规模庞大,基础模型并未在该任务中持续超越微调后的卷积神经网络。具体而言,EfficientNet-B0在大多数评估设置下,于受试者工作特征曲线下面积(AUC-ROC)和精确率/召回率曲线上排名第一或第二,而RETFound仅在OEFI数据集中表现突出。相反,FLAIR展现出具有竞争力的零样本性能,在适当提示下达到显著的AUC-PR得分。这些发现表明,即使经过微调,基础模型可能并不适合作为精细眼科任务如DME检测的工具,提示轻量级CNN在数据稀缺环境中仍是强有力的基线。
原文摘要 · Abstract (English)
Diabetic Macular Edema (DME) is a leading cause of vision loss among patients with Diabetic Retinopathy (DR). While deep learning has shown promising results for automatically detecting this condition from fundus images, its application remains challenging due the limited availability of annotated data. Foundation Models (FM) have emerged as an alternative solution. However, it is unclear if they can cope with DME detection in particular. In this paper, we systematically compare different FM and standard transfer learning approaches for this task. Specifically, we compare the two most popular FM for retinal images--RETFound and FLAIR--and an EfficientNet-B0 backbone, across different training regimes and evaluation settings in IDRiD, MESSIDOR-2 and OCT-and-Eye-Fundus-Images (OEFI). Results show that despite their scale, FM do not consistently outperform fine-tuned CNNs in this task. In particular, an EfficientNet-B0 ranked first or second in terms of area under the ROC and precision/recall curves in most evaluation settings, with RETFound only showing promising results in OEFI. FLAIR, on the other hand, demonstrated competitive zero-shot performance, achieving notable AUC-PR scores when prompted appropriately. These findings reveal that FM might not be a good tool for fine-grained ophthalmic tasks such as DME detection even after fine-tuning, suggesting that lightweight CNNs remain strong baselines in data-scarce environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。