arXiv:2502.06289eess.IVcs.AI2025-02中稿 · Ophthalmology Scie…

通用视觉模型在眼病检测上优于专用于视网膜的模型,但在系统性疾病预测上后者更优。

Is an Ultra Large Natural Image-Based Foundation Model Superior to a Retina-Specific Model for Detecting Ocular and Systemic Diseases?

  • 用通用大模型DINOv2和专有模型RETFound对比,分别在八组眼科数据集上测试
  • DINOv2-large在糖尿病视网膜病变检测中表现更好(AUROC 0.850-0.952)
  • RETFound在心脏病等系统性疾病预测中胜出(AUROC 0.732-0.796)

基础模型(FMs)正在改变医学领域。在眼科中,RETFound是基于140万张自然图像和160万张视网膜图像预训练的专用模型,展现出良好的临床适应性。而DINOv2是基于1.42亿张自然图像预训练的通用视觉模型,在非医疗领域表现良好,但其在临床任务中的适用性尚不明确。为此,我们通过微调RETFound和三个DINOv2模型(large、base、small),在八个标准化开源眼科数据集以及Moorfields AlzEye和UK Biobank数据集上,对眼病检测和系统性疾病预测任务进行了对比评估。结果显示,DINOv2-large在糖尿病视网膜病变检测中表现更优(AUROC=0.850–0.952 vs 0.823–0.944,三组数据集,P≤0.007),在多类眼病检测中也显著领先(AUROC=0.892 vs 0.846,P<0.001)。在青光眼检测中,DINOv2-base优于RETFound(AUROC=0.958 vs 0.940,P<0.001)。相反,RETFound在心力衰竭、心肌梗死和缺血性中风预测中均优于所有DINOv2模型(AUROC=0.732–0.796 vs 0.663–0.771,P<0.001)。这些趋势在仅使用10%微调数据时依然成立。研究揭示了通用与领域特定基础模型在不同任务中的优势场景,强调应根据任务需求选择合适的模型以优化临床表现。

原文摘要 · Abstract (English)

The advent of foundation models (FMs) is transforming medical domain. In ophthalmology, RETFound, a retina-specific FM pre-trained sequentially on 1.4 million natural images and 1.6 million retinal images, has demonstrated high adaptability across clinical applications. Conversely, DINOv2, a general-purpose vision FM pre-trained on 142 million natural images, has shown promise in non-medical domains. However, its applicability to clinical tasks remains underexplored. To address this, we conducted head-to-head evaluations by fine-tuning RETFound and three DINOv2 models (large, base, small) for ocular disease detection and systemic disease prediction tasks, across eight standardized open-source ocular datasets, as well as the Moorfields AlzEye and the UK Biobank datasets. DINOv2-large model outperformed RETFound in detecting diabetic retinopathy (AUROC=0.850-0.952 vs 0.823-0.944, across three datasets, all P<=0.007) and multi-class eye diseases (AUROC=0.892 vs. 0.846, P<0.001). In glaucoma, DINOv2-base model outperformed RETFound (AUROC=0.958 vs 0.940, P<0.001). Conversely, RETFound achieved superior performance over all DINOv2 models in predicting heart failure, myocardial infarction, and ischaemic stroke (AUROC=0.732-0.796 vs 0.663-0.771, all P<0.001). These trends persisted even with 10% of the fine-tuning data. These findings showcase the distinct scenarios where general-purpose and domain-specific FMs excel, highlighting the importance of aligning FM selection with task-specific requirements to optimise clinical performance.

基础模型眼科疾病模型对比医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。