研究生成模型在皮肤病图像中对不同肤色的公平性,发现模型对浅肤色表现更好。
Are generative models fair? A study of racial bias in dermatological image generation
- 用带感知损失的VAE生成和重建不同肤色皮肤图像
- 浅肤色图像生成质量更高,即使数据量相当
- 现有不确定性估计无法有效检测肤色偏差,适合医疗AI安全研究者
医学中的种族偏见,如皮肤病学领域,带来重大伦理与临床挑战。这通常源于训练数据中深色皮肤样本严重不足。尽管已有研究关注提升数据集多样性并缓解判别模型的不公平现象,但生成模型中的种族偏见仍缺乏深入探讨。本研究评估了生成模型在临床皮肤病学中的公平性。我们首先使用感知损失训练一个变分自编码器(VAE),以生成和重建跨不同肤色的高质量皮肤图像。基于Fitzpatrick17k数据集,分析种族偏见对模型表征与性能的影响。结果表明,模型性能随肤色代表性提升而改善;然而,即使在代表性相当的情况下,模型对浅肤色的生成效果仍更优。此外,VAE产生的不确定性估计无法有效反映其公平性。这些发现凸显了构建更具代表性的皮肤病数据集的必要性,也需深入理解偏见来源,并改进不确定性量化机制,以确保医疗生成模型的可信性。
原文摘要 · Abstract (English)
Racial bias in medicine, such as in dermatology, presents significant ethical and clinical challenges. This is likely to happen because there is a significant underrepresentation of darker skin tones in training datasets for machine learning models. While efforts to address bias in dermatology have focused on improving dataset diversity and mitigating disparities in discriminative models, the impact of racial bias on generative models remains underexplored. Generative models, such as Variational Autoencoders (VAEs), are increasingly used in healthcare applications, yet their fairness across diverse skin tones is currently not well understood. In this study, we evaluate the fairness of generative models in clinical dermatology with respect to racial bias. For this purpose, we first train a VAE with a perceptual loss to generate and reconstruct high-quality skin images across different skin tones. We utilize the Fitzpatrick17k dataset to examine how racial bias influences the representation and performance of these models. Our findings indicate that VAE performance is, as expected, influenced by representation, i.e. increased skin tone representation comes with increased performance on the given skin tone. However, we also observe, even independently of representation, that the VAE performs better for lighter skin tones. Additionally, the uncertainty estimates produced by the VAE are ineffective in assessing the model's fairness. These results highlight the need for more representative dermatological datasets, but also a need for better understanding the sources of bias in such model, as well as improved uncertainty quantification mechanisms to detect and address racial bias in generative models for trustworthy healthcare technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。