arXiv:2607.12987cs.CV2026-07中稿 · MICCAI 2026

用可控生成技术扩充皮肤病图像,提升诊断公平性与准确性。

Controllable Generation of Diverse Dermatological Imagery for Fair and Efficient Malignancy Classification

论文配图:Controllable Generation of Diverse Dermatological Imagery for Fair and Efficient Malignancy Classification
图 1 · 摘自论文原文
  • 通过合成健康皮肤与罕见病变迁移,实现跨肤色和部位的图像生成。
  • 仅需10张样本即可高效生成,合成训练下分类准确率达86.4%。
  • 适合关注医疗公平性、数据稀缺场景的研究者与开发者。

精准的皮肤科诊断需在不同人群间保持公平表现,但缺乏专家标注图像,尤其在低频肤色和罕见疾病上,阻碍了可衡量公平方法的发展。本文提出cgDDI(可控多样化皮肤病图像生成框架),支持:(1) 在不改变其他属性的前提下合成真实健康皮肤样本;(2) 非参数化地将单个罕见病变映射到新肤色与位置;(3) 仅需10个训练样本即可实现高效参数化生成。该框架兼容人工与自动分割掩码,可扩展至无预设病灶掩码的数据集。我们构建了一个包含656张图像的数据集,扩充超过400倍,并在两个数据集上验证:经活检确认的多样化皮肤病图像(DDI)与专家验证的Fitzpatrick17k(F17k)。在DDI基准上,纯合成训练下恶性肿瘤分类准确率达86.4%,结合真实数据微调后达90.9%的领先水平,且公平性指标优异。跨数据集实验显示,在未见过的F17k数据上准确率提升13.9%,尽管疾病重叠度极低。我们开源超过266,000张合成图像、代码与生成模型,以支持公平性研究,详见https://github.com/hectorcarrion/ControllableGenDDI。

原文摘要 · Abstract (English)

Accurate dermatological diagnosis naturally necessitates equitable performance across diverse populations, yet a systematic lack of expertly annotated images, especially for underrepresented skin tones and rare diseases, impedes progress toward measurably fair methods. We introduce cgDDI (Controllable Generation of Diverse Dermatological Imagery), a hybrid framework that (1) synthesizes realistic healthy skin samples without disturbing other input properties, (2) maps single-sample rare lesions onto novel skin-tones and locations non-parametrically, and (3) allows for efficient parametric generation with as few as 10 training samples. The framework supports both human and automated segmentation masking, enabling scalability to datasets without pre-made lesion masks. We grow a 656-image dataset by more than 400x and validate across two datasets: biopsy-confirmed Diverse Dermatology Images (DDI) and expert-verified Fitzpatrick17k (F17k). On the DDI benchmark, we achieve malignancy classification accuracy of 86.4% under synthetic-only training and 90.9% state-of-the-art performance with real data fine-tuning, alongside leading fairness metrics. Cross-dataset experiments show +13.9% accuracy improvements on unseen F17k data despite minimal disease overlap. We openly release 266k+ synthetic images, code, and generative models to further support fairness research at https://github.com/hectorcarrion/ControllableGenDDI.

皮肤病生成图像合成公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。