arXiv:2602.10624cs.CVcs.AI2026-02被引 3

无需微调的皮肤病多模态模型,提升医生诊断准确率

A Vision-Language Foundation Model for Zero-shot Clinical Collaboration and Automated Concept Discovery in Dermatology

论文配图:A Vision-Language Foundation Model for Zero-shot Clinical Collaboration and Automated Concept Discovery in Dermatology
图 1 · 摘自论文原文
  • 用400万张多模态数据训练,通过掩码潜空间与对比学习构建模型
  • 零样本诊断在20个基准上达顶尖水平,基层医生诊断准确率翻倍
  • 可解释潜空间发现临床概念,抑制伪影偏差,适合临床协作使用

医疗基础模型在受控基准中表现良好,但广泛应用仍受限于任务特定微调。本文提出DermFM-Zero,一种基于掩码潜空间建模和对比学习,在超过400万个多模态数据点上训练的皮肤科视觉-语言基础模型。我们在20个涵盖零样本诊断与多模态检索的基准上评估该模型,未进行任务特定适配即达到领先性能。进一步在涉及1100多名临床医生的三个跨国读者研究中验证其零样本能力:在初级保健场景下,AI辅助使全科医生对98种皮肤疾病的鉴别诊断准确率接近翻倍;在专科场景中,模型在多模态皮肤癌评估中显著优于持证皮肤科医生;在协作工作流中,非专家在AI协助下超越未辅助的专家,并提升管理恰当性。最后,我们发现DermFM-Zero的潜空间具有可解释性:稀疏自编码器无监督地解耦出临床有意义的概念,优于预定义词汇方法,且能针对性抑制伪影引发的偏差,增强鲁棒性而无需重新训练。这些结果表明,基础模型可在无需微调的情况下提供有效、安全、透明的零样本临床决策支持。

原文摘要 · Abstract (English)

Medical foundation models have shown promise in controlled benchmarks, yet widespread deployment remains hindered by reliance on task-specific fine-tuning. Here, we introduce DermFM-Zero, a dermatology vision-language foundation model trained via masked latent modelling and contrastive learning on over 4 million multimodal data points. We evaluated DermFM-Zero across 20 benchmarks spanning zero-shot diagnosis and multimodal retrieval, achieving state-of-the-art performance without task-specific adaptation. We further evaluated its zero-shot capabilities in three multinational reader studies involving over 1,100 clinicians. In primary care settings, AI assistance enabled general practitioners to nearly double their differential diagnostic accuracy across 98 skin conditions. In specialist settings, the model significantly outperformed board-certified dermatologists in multimodal skin cancer assessment. In collaborative workflows, AI assistance enabled non-experts to surpass unassisted experts while improving management appropriateness. Finally, we show that DermFM-Zero's latent representations are interpretable: sparse autoencoders unsupervisedly disentangle clinically meaningful concepts that outperform predefined-vocabulary approaches and enable targeted suppression of artifact-induced biases, enhancing robustness without retraining. These findings demonstrate that a foundation model can provide effective, safe, and transparent zero-shot clinical decision support.

皮肤病多模态零样本可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。