arXiv:2505.05291eess.IVcs.AI2025-05被引 2

自然图像预训练的视觉模型在眼底病检测中表现更优,挑战了领域内预训练必要性。

Benchmarking Ophthalmology Foundation Models for Clinically Significant Age Macular Degeneration Detection

  • 用六种自监督预训练ViT在七大数据集上对比性能
  • 自然图像预训练模型AUROC达0.80-0.97,优于领域模型
  • 适用于眼科医生和算法研究者参考模型选择

自监督学习(SSL)使视觉变换器(ViTs)能从大规模自然图像中学习鲁棒表征,提升跨域泛化能力。在眼底成像领域,基于自然或眼科数据预训练的基础模型已展现潜力,但领域内预训练的优势尚不明确。我们对六种SSL预训练的ViT在七个数字眼底图像(DFI)数据集上进行基准测试,总样本量达7万张专家标注图像,任务为中重度年龄相关性黄斑变性(AMD)识别。结果显示,基于自然图像预训练的iBOT模型在分布外泛化能力最强,AUROC达0.80–0.97,优于领域特定模型(AUROC 0.78–0.96)及无预训练的ViT-L(AUROC 0.68–0.91)。该结果凸显基础模型在提升AMD识别中的价值,并质疑了领域内预训练的必要性。此外,我们发布了BRAMD数据集(n=587),包含巴西来源的带AMD标签的DFI,为开源数据。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has enabled Vision Transformers (ViTs) to learn robust representations from large-scale natural image datasets, enhancing their generalization across domains. In retinal imaging, foundation models pretrained on either natural or ophthalmic data have shown promise, but the benefits of in-domain pretraining remain uncertain. To investigate this, we benchmark six SSL-pretrained ViTs on seven digital fundus image (DFI) datasets totaling 70,000 expert-annotated images for the task of moderate-to-late age-related macular degeneration (AMD) identification. Our results show that iBOT pretrained on natural images achieves the highest out-of-distribution generalization, with AUROCs of 0.80-0.97, outperforming domain-specific models, which achieved AUROCs of 0.78-0.96 and a baseline ViT-L with no pretraining, which achieved AUROCs of 0.68-0.91. These findings highlight the value of foundation models in improving AMD identification and challenge the assumption that in-domain pretraining is necessary. Furthermore, we release BRAMD, an open-access dataset (n=587) of DFIs with AMD labels from Brazil.

眼底图像基础模型AMD检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。