arXiv:2509.03421eess.IVcs.CV2025-09被引 7

专用眼底模型比通用视觉模型在眼科疾病检测上更优,数据效率更高。

Generalist versus Specialist Vision Foundation Models for Ocular Disease and Oculomics

  • 对比通用模型与专用眼底模型的适应能力,采用微调和线性探测两种策略。
  • 专用模型RETFound-DINOv2在眼病检测和系统性疾病预测中表现更优,数据效率更高。
  • 适合临床医学应用,尤其关注眼底影像分析的研究者与开发者。

医学基础模型通过大规模临床数据预训练,在多种临床相关任务中表现出色。RETFound基于近百万张眼底图像训练,代表了该领域的进展。然而,随着DINOv2和DINOv3等更强大、规模更大的通用基础模型出现,人们开始质疑领域专用预训练是否仍必要,以及现存差距为何。为此,我们系统评估了DINOv2和DINOv3在眼底图像任务中的可适应性,与两个专用模型RETFound-MAE和RETFound-DINOv2进行对比。采用微调和线性探测两种策略,评估其在眼病检测与系统性疾病预测上的表现。进一步分析数据效率与适配效率,揭示性能与计算成本之间的权衡。结果表明,尽管通用模型在多任务适应上表现强劲,但RETFound-DINOv2在眼病检测与眼组学任务中始终优于通用模型,展现出更强的泛化能力和数据效率。研究提示:专用眼底基础模型仍是临床应用的最优选择;而通用模型与专用模型间的差距缩小,表明持续的数据与模型扩展仍能带来领域相关收益,有望成为未来医学基础模型的强大基石。

原文摘要 · Abstract (English)

Medical foundation models, pre-trained with large-scale clinical data, demonstrate strong performance in diverse clinically relevant applications. RETFound, trained on nearly one million retinal images, exemplifies this approach in applications with retinal images. However, the emergence of increasingly powerful and multifold larger generalist foundation models such as DINOv2 and DINOv3 raises the question of whether domain-specific pre-training remains essential, and if so, what gap persists. To investigate this, we systematically evaluated the adaptability of DINOv2 and DINOv3 in retinal image applications, compared to two specialist RETFound models, RETFound-MAE and RETFound-DINOv2. We assessed performance on ocular disease detection and systemic disease prediction using two adaptation strategies: fine-tuning and linear probing. Data efficiency and adaptation efficiency were further analysed to characterise trade-offs between predictive performance and computational cost. Our results show that although scaling generalist models yields strong adaptability across diverse tasks, RETFound-DINOv2 consistently outperforms these generalist foundation models in ocular-disease detection and oculomics tasks, demonstrating stronger generalisability and data efficiency. These findings suggest that specialist retinal foundation models remain the most effective choice for clinical applications, while the narrowing gap with generalist foundation models suggests that continued data and model scaling can deliver domain-relevant gains and position them as strong foundations for future medical foundation models.

眼底影像基础模型医学AI数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。