小模型比大模型更适合多数眼科影像分类任务
When Do Domain-Specific Foundation Models Justify Their Cost? A Systematic Evaluation Across Retinal Imaging Tasks
- 用小型通用模型替代大型专业模型,效果更优且节省算力
- 在多数任务中,小型模型准确率接近顶尖水平,如糖尿病黄斑水肿达99.24%
- 仅在复杂分级任务中(如糖尿病视网膜病变)才需专用大模型
大型视觉基础模型虽被广泛用于视网膜疾病分类,但缺乏系统证据说明其参数规模的必要性。本文通过四个视网膜影像分类任务(OCT 8类、DME 3类、DR 5类、GL 3类)评估12-13种模型配置,涵盖视觉变换器(22.8M-86.6M参数)、Swin Transformer(27.6M-28.3M)、ConvNeXt(28.6M)及领域专用模型RETFound(303M),在相同训练条件下对比初始化策略。结果表明:预训练普遍提升性能(5.18%-18.41%),且随任务难度增加而增强;小型架构(27-29M)在帕累托前沿占优,SwinV2-tiny在三个数据集上表现最佳;仅在挑战性的DR分级任务中(准确率71.15%),303M的RETFound才值得其计算开销,其余任务中ImageNet预训练已足够,如DME准确率达99.24%,OCT为97.96%。彩色眼底摄影任务预训练增益(9.13%-18.41%)高于OCT(5.18%)。因此,多数视网膜分类任务可用紧凑通用模型实现近最优性能,专用大模型仅适用于极端类别不平衡下的精细区分。
原文摘要 · Abstract (English)
Large vision foundation models have been widely adopted for retinal disease classification without systematic evidence justifying their parameter requirements. In the present work we address two critical questions: First, are large domain-specific foundation models essential, or do compact general-purpose architectures suffice? Second, does specialized retinal pretraining justify its computational cost? To answer this, we benchmark initialization strategies across four retinal imaging classification tasks spanning Optical Coherence Tomography (OCT) and Color Fundus Photography (CFP) modalities: 8-class OCT classification, 3-class diabetic macular edema (DME), 5-class diabetic retinopathy (DR), and 3-class glaucoma (GL) detection. We evaluate 12-13 model configurations per task, including vision transformers (22.8M-86.6M parameters), Swin Transformers (27.6M-28.3M), ConvNeXt (28.6M), and the domain-specific RETFound models (303M), under identical training conditions. Our results challenge prevailing assumptions: First, we demonstrate that pretraining provides universal benefits (5.18-18.41% improvement), scaling with task difficulty. Second, compact architectures (27-29M) dominate Pareto frontiers; SwinV2-tiny achieves top-1 performance on three datasets. Third, RETFound (303M) justifies its computational cost only for challenging DR grading (accuracy of 71.15%), while ImageNet pretraining proves to be sufficient with all other tasks (DME accuracy: 99.24%, OCT accuracy: 97.96%). CFP tasks show larger pretraining accuracy gains (9.13-18.41%) than OCT (5.18%). Thus, the evidence suggests that compact general-purpose models deliver near-optimal performance for most retinal classification tasks; specialized foundation models warranted only for fine-grained discrimination under extreme class imbalance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。