arXiv:2508.15986cs.CVcs.AI2025-08被引 3

用百万张合成眼底图训练多病种分类模型,效果媲美真实数据。

Automated Multi-label Classification of Eleven Retinal Diseases: A Benchmark of Modern Architectures and a Meta-Ensemble on a Large Synthetic Dataset

  • 用六种先进模型在合成数据上端到端训练,多标签分类。
  • 集成模型在内部验证集上宏AUC达0.9973,跨数据集泛化能力好。
  • 为合成数据训练眼科AI提供基准,适合想用合成数据的研究者。

多标签深度学习模型在视网膜疾病分类中的发展常受限于高质量临床标注数据的稀缺性,这源于患者隐私顾虑和高昂成本。最近发布的高保真合成数据集SynFundus-1M包含超过一百万张眼底图像,为突破这一瓶颈提供了新机遇。为建立该数据资源的性能基准,我们构建了端到端深度学习流程,采用五折多标签分层交叉验证策略,训练六种现代架构(ConvNeXtV2、SwinV2、ViT、ResNet、EfficientNetV2及RETFound基础模型)对十一类视网膜疾病进行分类。进一步通过堆叠各模型的留出预测结果,并使用XGBoost构建元集成模型。最终集成模型在内部验证集上表现最佳,宏平均受试者工作特征曲线下面积(macro-AUC)达到0.9973。关键的是,模型在三个不同真实临床数据集上表现出强泛化能力:在合并的糖尿病视网膜病变数据集上AUC为0.7972,在AIROGS青光眼数据集上AUC为0.9126,在多标签RFMiD数据集上宏AUC为0.8800。本研究为大规模合成数据上的研究提供了稳健基线,证明仅在合成数据上训练的模型也能准确识别多种病变,并有效泛化至真实临床图像,为推动眼科全面AI系统的发展提供了可行路径。

原文摘要 · Abstract (English)

The development of multi-label deep learning models for retinal disease classification is often hindered by the scarcity of large, expertly annotated clinical datasets due to patient privacy concerns and high costs. The recent release of SynFundus-1M, a high-fidelity synthetic dataset with over one million fundus images, presents a novel opportunity to overcome these barriers. To establish a foundational performance benchmark for this new resource, we developed an end-to-end deep learning pipeline, training six modern architectures (ConvNeXtV2, SwinV2, ViT, ResNet, EfficientNetV2, and the RETFound foundation model) to classify eleven retinal diseases using a 5-fold multi-label stratified cross-validation strategy. We further developed a meta-ensemble model by stacking the out-of-fold predictions with an XGBoost classifier. Our final ensemble model achieved the highest performance on the internal validation set, with a macro-average Area Under the Receiver Operating Characteristic Curve (AUC) of 0.9973. Critically, the models demonstrated strong generalization to three diverse, real-world clinical datasets, achieving an AUC of 0.7972 on a combined DR dataset, an AUC of 0.9126 on the AIROGS glaucoma dataset and a macro-AUC of 0.8800 on the multi-label RFMiD dataset. This work provides a robust baseline for future research on large-scale synthetic datasets and establishes that models trained exclusively on synthetic data can accurately classify multiple pathologies and generalize effectively to real clinical images, offering a viable pathway to accelerate the development of comprehensive AI systems in ophthalmology.

眼科AI合成数据多标签分类眼底图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。