arXiv:2409.07422eess.IVcs.CV2024-09被引 4

用可控生成提升糖尿病视网膜病变诊断,解决标注数据少难题

Controllable retinal image synthesis using conditional StyleGAN and latent space manipulation for improved diagnosis and grading of diabetic retinopathy

  • 通过条件StyleGAN实现病程与视觉特征的精准控制
  • 生成图像使分类器准确率达98.09%,分级任务达到83.33%准确率
  • 无需额外网络,适合医学图像增强与模型训练场景

糖尿病视网膜病变(DR)是糖尿病导致的视网膜血管损伤,及时检测可降低失明风险。但严重病例标注数据稀缺,制约了分级模型训练。本文提出一种可控生成高保真、多样化DR眼底图像的框架,显著提升分类性能。仅使用条件StyleGAN,即可在不依赖特征掩码或辅助网络的前提下,对病变程度及视盘、血管结构、病灶区域等视觉特征实现全面控制。基于SeFa算法识别潜在语义,对条件生成图像进行潜空间操控,进一步增强数据多样性。此外,提出一种新颖有效的基于SeFa的数据增强策略,帮助分类器聚焦判别性区域,忽略冗余信息。在APTOS 2019数据集上的实验表明,使用该方法训练的ResNet50模型在DR检测中达98.09%准确率、99.44%特异性、99.45%精确率和98.09%F1分数;将合成图像用于分级训练,准确率达83.33%,二次加权卡帕系数为87.64%,特异性95.67%,精确率72.24%。生成图像具有极佳真实感,分类性能优于近期研究。

原文摘要 · Abstract (English)

Diabetic retinopathy (DR) is a consequence of diabetes mellitus characterized by vascular damage within the retinal tissue. Timely detection is paramount to mitigate the risk of vision loss. However, training robust grading models is hindered by a shortage of annotated data, particularly for severe cases. This paper proposes a framework for controllably generating high-fidelity and diverse DR fundus images, thereby improving classifier performance in DR grading and detection. We achieve comprehensive control over DR severity and visual features (optic disc, vessel structure, lesion areas) within generated images solely through a conditional StyleGAN, eliminating the need for feature masks or auxiliary networks. Specifically, leveraging the SeFa algorithm to identify meaningful semantics within the latent space, we manipulate the DR images generated conditionally on grades, further enhancing the dataset diversity. Additionally, we propose a novel, effective SeFa-based data augmentation strategy, helping the classifier focus on discriminative regions while ignoring redundant features. Using this approach, a ResNet50 model trained for DR detection achieves 98.09% accuracy, 99.44% specificity, 99.45% precision, and an F1-score of 98.09%. Moreover, incorporating synthetic images generated by conditional StyleGAN into ResNet50 training for DR grading yields 83.33% accuracy, a quadratic kappa score of 87.64%, 95.67% specificity, and 72.24% precision. Extensive experiments conducted on the APTOS 2019 dataset demonstrate the exceptional realism of the generated images and the superior performance of our classifier compared to recent studies.

医学图像生成风格迁移数据增强糖尿病视网膜病变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。