arXiv:2609.02111cs.CVcs.AI2026-09

皮肤色度与疾病分布如何影响皮肤病AI模型泛化?

Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization Gap

  • 区分皮肤色度与疾病分布对模型泛化的相对影响
  • 疾病分布偏移导致性能下降,远超肤色差异的影响
  • 皮肤病预训练模型仅需少量标注样本即可高效适配新场景

皮肤病人工智能模型主要在浅肤色、以癌症为主的图像数据集上训练,但常被提议部署于资源匮乏环境中,这些环境中的患者在肤色和疾病分布上均与训练数据不同。我们探究模型泛化不佳的主要原因是肤色代表性不足还是疾病分布偏移。评估了癌症训练基线(基于HAM10000和ISIC 2019微调的ResNet-50)、两个皮肤病基础模型(DermLIP和MONET)以及作为冻结特征提取器的一般视觉模型(DINOv3)。在分肤色匹配疾病的数据集(DDI)和疾病偏移但肤色多样的数据集(SCIN)上进行测试。结果显示,在所考察场景中,疾病分布偏移的影响大于肤色差异:癌症基线在迁移到不熟悉临床条件时,平衡准确率从0.62降至0.21;而同病种内肤色差距较小(0.10–0.18),且不一致。无标签表示分析表明,这种失败反映的是表征能力限制而非仅缺少输出标签:癌症专精特征在陌生条件下聚类效果差(kNN纯度提升+0.06,仅略高于随机),而皮肤病预训练特征保留更强可迁移结构(+0.23)。最后,我们证明表示质量可预测轻量适配下的可恢复性能:从皮肤病基础模型出发,每类约10个标注样本即可恢复大部分可达到性能。我们公开评估协议和代码,支持可复现的皮肤病AI泛化审计。

原文摘要 · Abstract (English)

Dermatology artificial intelligence (AI) models are predominantly trained on light-skinned, cancer-focused image collections, yet they are increasingly proposed for deployment in resource-constrained settings where patients differ from training populations along two confounded axes: skin tone and disease distribution. We investigate whether poor generalization is primarily caused by skin-tone underrepresentation or disease-distribution shift. We evaluate a cancer-trained baseline (ResNet-50 fine-tuned on HAM10000 and ISIC 2019), two dermatology foundation models (DermLIP and MONET), and a general-purpose vision model (DINOv3) as frozen feature extractors. Models are evaluated on a tone-stratified disease-matched dataset (Diverse Dermatology Images, DDI) and a disease-shifted tone-diverse dataset (Skin Condition Image Network, SCIN). Our results show that disease-distribution shift contributes more than skin tone in the evaluated settings. The cancer baseline decreases from 0.62 to 0.21 balanced accuracy when transferred to unfamiliar clinical conditions, while the within-disease skin-tone gap is smaller (0.10-0.18) and inconsistent. Label-free representation analysis shows that this failure reflects a representational limitation rather than only missing output labels: cancer-specialized features poorly cluster unfamiliar conditions (kNN purity lift +0.06 over chance), whereas dermatology-pretrained features retain stronger transferable structure (+0.23). Finally, we show that representation quality predicts recoverable performance under lightweight adaptation. Starting from dermatology foundation models, approximately ten labeled examples per clinical category recover most attainable performance. We release the evaluation protocol and code to support reproducible auditing of dermatology AI generalization.

皮肤病AI模型泛化公平性少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。