用分层特征生成高保真眼底图像,解决医学数据少的难题。
FundusGAN: A Hierarchical Feature-Aware Generative Framework for High-Fidelity Fundus Image Generation
- 通过金字塔结构提取多尺度眼底特征,保留大结构和细微病灶。
- 在DDR数据集上SSIM达0.8863,FID为54.2,生成图像质量领先。
- 生成图像可提升诊断准确率最高6.49%,适合数据受限的医学研究。
近年来,像RetFound这样的眼科基础模型展现出强大诊断能力,但需要海量数据预训练,极大阻碍了其发展与应用。为解决这一关键挑战,我们提出FundusGAN,一种专为高保真眼底图像合成设计的分层特征感知生成框架。该方法在编码器中引入特征金字塔网络,全面提取多尺度信息,同时捕捉大范围解剖结构与细微病理特征。生成器采用改进的StyleGAN架构,结合空洞卷积与优化上采样策略,有效保持视网膜关键结构并增强病灶细节表现。在DDR、DRIVE和IDRiD数据集上的综合评估显示,FundusGAN在多个指标上持续优于现有方法(于DDR上SSIM: 0.8863,FID: 54.2,KID: 0.0436)。此外,疾病分类实验表明,使用FundusGAN生成图像扩充训练数据后,多种CNN模型诊断准确率显著提升(以ResNet50为例最高达6.49%)。这些结果证明FundusGAN是应对眼科人工智能研究中数据稀缺问题的有力工具,有助于构建更鲁棒、泛化性更强的诊断系统,减少对大规模临床数据采集的依赖。
原文摘要 · Abstract (English)
Recent advancements in ophthalmology foundation models such as RetFound have demonstrated remarkable diagnostic capabilities but require massive datasets for effective pre-training, creating significant barriers for development and deployment. To address this critical challenge, we propose FundusGAN, a novel hierarchical feature-aware generative framework specifically designed for high-fidelity fundus image synthesis. Our approach leverages a Feature Pyramid Network within its encoder to comprehensively extract multi-scale information, capturing both large anatomical structures and subtle pathological features. The framework incorporates a modified StyleGAN-based generator with dilated convolutions and strategic upsampling adjustments to preserve critical retinal structures while enhancing pathological detail representation. Comprehensive evaluations on the DDR, DRIVE, and IDRiD datasets demonstrate that FundusGAN consistently outperforms state-of-the-art methods across multiple metrics (SSIM: 0.8863, FID: 54.2, KID: 0.0436 on DDR). Furthermore, disease classification experiments reveal that augmenting training data with FundusGAN-generated images significantly improves diagnostic accuracy across multiple CNN architectures (up to 6.49\% improvement with ResNet50). These results establish FundusGAN as a valuable foundation model component that effectively addresses data scarcity challenges in ophthalmological AI research, enabling more robust and generalizable diagnostic systems while reducing dependency on large-scale clinical data collection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。