用定制模型提升皮肤疾病诊断公平性,解决大模型对深肤色人群的识别偏差。
Adapting Large Language Models to Mitigate Skin Tone Biases in Clinical Dermatology Tasks: A Mixed-Methods Study
- 基于SkinGPT-4微调新模型,针对不同肤色优化皮肤病分类任务。
- 定制模型在6种疾病上平均准确率达0.75,肤色间公平性差距缩小至0.21。
- 适合医疗AI开发者与皮肤科医生,关注公平性与临床实用性的研究者。
SkinGPT-4是一种大型视觉语言模型,利用标注的皮肤疾病图像增强弱势群体的临床工作流程。然而其训练数据主要代表浅肤色,导致对深肤色患者的诊断准确性下降。本研究基于开源的SCIN数据集,评估了SkinGPT-4在常见皮肤病(如湿疹、过敏性接触性皮炎、银屑病)中按肤色分层的表现偏差。采用SkinGPT-4骨干网络开发微调模型以适应自定义皮肤疾病分类任务,并探索偏见缓解策略。由认证皮肤科医生对300个SCIN病例进行临床评估,涵盖诊断准确率、信息量、医生实用性与患者实用性。模型公平性指标(包括人口均等性和等化几率)在不同肤色间计算。SkinGPT-4在菲茨帕特里克分型中平均人口均等性为0.10,轻肤色与深肤色间差异达0.10–0.15。模型幻觉(如解剖结构与伪影误判)发生率为17.8%。定制模型在视觉相似疾病对上平均F1、精确率与AUROC分别为0.75、0.78和0.78。公平性分析显示平均人口均等性达0.75,最大肤色差异为0.21。最优模型在菲茨帕特里克I-VI型中分别达到0.83、0.83、0.76、0.89、0.90、0.90的均等性分数,表明其具备强鲁棒性。结果表明,使用现有骨干模型可有效训练出兼具准确性和公平性的定制化皮肤疾病分类系统。
原文摘要 · Abstract (English)
SkinGPT-4, a large vision-language model, leverages annotated skin disease images to augment clinical workflows in underserved communities. However, its training dataset predominantly represents lighter skin tones, limiting diagnostic accuracy for darker tones. Here, we evaluated performance biases in SkinGPT-4 across skin tones on common skin diseases, including eczema, allergic-contact dermatitis, and psoriasis using the open-sourced SCIN dataset. We leveraged the SkinGPT-4 backbone to develop finetuned models for custom skin disease classification tasks and explored bias mitigation strategies. Clinical evaluation by board-certified dermatologists on six relevant skin diseases from 300 SCIN cases assessed images for diagnostic accuracy, informativity, physician utility, and patient utility. Model fairness metrics, including demographic parity and equalized odds, were calculated across skin tones. SkinGPT-4 achieved an average demographic parity of 0.10 across Fitzpatrick types, with notable differences of 0.10-0.15 between lightest and darkest tones across evaluation metrics. Model hallucinations in artifacts and anatomy occurred at a rate of 17.8. Our customized models achieved average F1, precision, and AUROC of 0.75, 0.78, and 0.78 across visually similar disease pairs. Fairness analysis showed an average demographic parity of 0.75, with a maximum disparity of 0.21 across skin tones. The best model achieved parity scores of 0.83, 0.83, 0.76, 0.89, 0.90, and 0.90 for Fitzpatrick I-VI, indicating robust fairness. Large language models such as SkinGPT-4 showed weaker performance on darker tones. Model biases exist across evaluation criteria, and hallucinations may affect diagnostic efficacy. These findings demonstrate the efficacy of training accurate, fair models using existing backbones for custom skin disease classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。