SkinGPT-R1让皮肤科AI诊断更公平透明,跨肤色表现稳定。
Trustworthy and Fair SkinGPT-R1 for Democratizing Dermatological Reasoning across Diverse Ethnicities
- 用思维链+专家混合模型,生成可解释的诊断报告。
- 在7个数据集上6项领先,长尾分类准确率达82.50%。
- 显著降低肤色偏差,低性能群体提升5倍,适合临床辅助。
皮肤病学AI的临床应用受限于推理不透明和不同肤色间的性能差异。本文提出SkinGPT-R1,一个融合思维链诊断推理与公平感知的专家混合架构的多模态大语言模型,实现可解释且公平的皮肤疾病诊断。通过冻结推理主干的参数高效适配,SkinGPT-R1生成包含视觉发现、鉴别推理与最终诊断的结构化报告。在涵盖多种病种与成像条件的七个外部数据集上,其在六项基准测试中达到最优表现,包括在40类长尾分类任务上达82.50%准确率(较领先基线提升19.30%)。五位持证皮肤科医生对1000例表型均衡病例进行盲评,平均得分为3.6/5,安全性(3.8)与推理连贯性(3.6)评分最高,表明生成推理具备临床安全性、逻辑一致性,适合辅助诊断决策。关键在于,SkinGPT-R1在全菲茨帕特里克肤色谱中有效缓解算法偏见,在Fitz17k基准上最差组表现达41.40%,并在DDI数据集上使低性能组准确率相对提升五倍,优于标准多模态基线。这些结果建立了可信、公平、可解释的AI辅助皮肤科诊断新范式。
原文摘要 · Abstract (English)
The clinical translation of dermatological AI is hindered by opaque reasoning and systematic performance disparities across skin tones. Here we present SkinGPT-R1, a multimodal large language model that integrates chain-of-thought diagnostic reasoning with a fairness-aware mixture-of-experts architecture for interpretable and equitable skin disease diagnosis. Through parameter-efficient adaptation of a frozen reasoning backbone, SkinGPT-R1 generates structured diagnostic reports comprising visual findings, differential reasoning, and final diagnosis. Across seven external datasets spanning diverse pathologies and imaging conditions, SkinGPT-R1 achieves state-of-the-art accuracy on six benchmarks, including 82.50\% on a challenging 40-class long-tail classification task (+19.30\% over leading baselines). Blinded evaluation by five board-certified dermatologists on 1,000 phenotypically balanced cases yields a mean score of 3.6 out of 5, with the highest ratings in safety (3.8) and reasoning coherence (3.6), indicating that the generated rationales are clinically safe, logically grounded, and suitable for supporting diagnostic decision-making. Critically, SkinGPT-R1 mitigates algorithmic bias across the full Fitzpatrick spectrum, achieving a robust worst-group performance of 41.40\% on the Fitz17k benchmark and a five-fold relative improvement in lower-bound accuracy on the DDI dataset compared to standard multimodal baselines. These results establish a framework for trustworthy, fair, and explainable AI-assisted dermatological diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。