arXiv:2512.17585eess.IVcs.CV2025-12

对比生成模型与预处理对皮肤癌图像合成的影响,发现模型选择比预处理更重要。

SkinGenBench: Generative Model and Preprocessing Effects for Synthetic Dermoscopic Augmentation in Melanoma Diagnosis

  • 比较StyleGAN2-ADA与扩散模型在合成皮肤镜图像中的表现。
  • 生成数据使黑色素瘤检测F1分数提升8%~15%,最高达0.88。
  • 复杂预处理可能削弱关键纹理特征,影响诊断效果。

本文提出SkinGenBench,一个系统性的生物医学成像基准,研究预处理复杂度与生成模型选择在合成皮肤镜图像增强及下游黑色素瘤诊断中的交互作用。基于HAM10000和MILK10K数据集的14,116张皮肤镜图像,涵盖五类病灶,评估两种代表性生成范式:StyleGAN2-ADA与去噪扩散概率模型(DDPM)。在基础几何增强与高级伪影去除流程下,通过FID、KID、IS等感知与分布指标、特征空间分析,以及五种下游分类器的诊断性能,评估合成黑色素瘤图像。实验表明,生成架构选择对图像保真度与诊断效用的影响强于预处理复杂度。StyleGAN2-ADA生成图像更贴近真实分布,FID≈65.5,KID≈0.05;扩散模型样本方差更高但感知保真度下降且类别锚定弱。高级伪影去除仅带来微弱指标提升,下游诊断增益有限,提示可能抑制临床相关纹理线索。相比之下,合成数据增强显著提升黑色素瘤检测性能,F1分数绝对提升8%~15%,ViT-B/16达到F1≈0.88,ROC-AUC≈0.98,较无增强基线提升约14%。

原文摘要 · Abstract (English)

This work introduces SkinGenBench, a systematic biomedical imaging benchmark that investigates how preprocessing complexity interacts with generative model choice for synthetic dermoscopic image augmentation and downstream melanoma diagnosis. Using a curated dataset of $14,116$ dermoscopic images from HAM10000 and MILK10K across five lesion classes, we evaluate the two representative generative paradigms: StyleGAN2-ADA and Denoising Diffusion Probabilistic Models (DDPMs) under basic geometric augmentation and advanced artifact removal pipelines. Synthetic melanoma images are assessed using established perceptual and distributional metrics (FID, KID, IS), feature space analysis, and their impact on diagnostic performance across five downstream classifiers. Experimental results demonstrate that generative architecture choice has a stronger influence on both image fidelity and diagnostic utility than preprocessing complexity. StyleGAN2-ADA consistently produced synthetic images more closely aligned with real data distributions, achieving the lowest FID ($\approx 65.5$) and KID ($\approx 0.05$), while diffusion models generated higher variance samples at the cost of reduced perceptual fidelity and class anchoring. Advanced artifact removal yielded only marginal improvements in generative metrics and provided limited downstream diagnostic gains, suggesting possible suppression of clinically relevant texture cues. In contrast, synthetic data augmentation substantially improved melanoma detection with $8$-$15$\% absolute gains in melanoma F1-score, and ViT-B/16 achieving F1 $\approx 0.88$ and ROC-AUC $\approx 0.98$, representing an improvement of approximately $14\%$ over non-augmented baselines. Our code can be found at https://github.com/adarsh-crafts/SkinGenBench

图像生成皮肤癌数据增强生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。