对比四类模型在多疾病眼底筛查中的表现,发现注意力模型更优。
Benchmarking Convolutional, Transformer, Hybrid, and Vision Language Models for Multi Disease Retinal Screening
- 系统比较卷积、Transformer、混合与视觉语言模型在眼底图像上的表现。
- 注意力模型在二分类任务中AUC超84%,混合模型在多标签任务中F1最优。
- 结果为临床部署提供可复现的模型选型参考,适合医疗AI研究者。
现代深度学习为自动化眼底筛查提供了强大工具,但不同视觉模型家族在真实多疾病场景和域偏移下的表现仍不清晰。本文在Retinal Fundus Multi-disease Image Dataset (RFMiD) 上基准测试了十二种架构,涵盖四类模型:卷积神经网络、视觉Transformer、CNN-Transformer混合结构、视觉语言模型。评估两项任务:任意视网膜疾病二分类筛查,以及28类疾病的多标签分类。采用标准化训练、校准与评估协议,报告在特异性接近80%的临床相关操作点下的AUC、F1、精确率、召回率和敏感性。在RFMiD上,所有模型在二分类任务中表现良好(AUC > 84%),注意力模型表现最佳。SwinTiny及混合模型CoAtNet0、MaxViTTiny在二分类中表现最强,并在多标签任务中提升宏平均与微平均F1。视觉语言模型(如CLIP ViT-B/16、SigLIP-Base384)性能与CNN基线相当,但未超越最佳Transformer与混合骨干模型。在外部验证集Messidor-2上对可转诊糖尿病视网膜病变的评估中,AUC范围为66.8%至84.7%,再次显示混合与Transformer模型优势。本研究为多疾病眼底筛查中的模型选择提供可复现参考,指导未来自动化筛查系统的临床部署。
原文摘要 · Abstract (English)
Modern deep learning offers powerful tools for automated retinal screening, but it remains unclear how different visual model families compare in realistic multi-disease settings and under domain shift. In this work, we benchmark twelve architectures across four model families: convolutional neural networks, vision transformers, hybrid CNN-transformer backbones, and vision-language models, using the Retinal Fundus Multi-disease Image Dataset (RFMiD). We evaluate two tasks: binary screening for any retinal disease and multi-label classification across 28 disease classes. Using standardized training, calibration, and evaluation protocols, we report AUC, F1, precision, recall, and sensitivity at a clinically relevant operating point with specificity near 80%. On RFMiD, all architectures perform well on binary screening, with AUC above 84%, but attention-based models perform best. SwinTiny and the hybrid CoAtNet0 and MaxViTTiny models achieve the strongest binary screening results and improve macro and micro F1 in the multi-label setting. Vision-language models, including CLIP ViT-B/16 and SigLIP-Base384, are competitive with CNN baselines but do not surpass the best transformer and hybrid backbones. In external validation on Messidor-2 for referable diabetic retinopathy, AUC ranges from 66.8% to 84.7%, with hybrid and transformer models again showing strong performance. These results provide a reproducible reference for model selection in multi-disease retinal screening and guide future automated screening tools for clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。