构建首个百万级皮肤病视觉语言数据集,助力医学AI精准诊断
Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology

- 基于临床术语体系构建跨模态图文对,覆盖390种皮肤疾病
- 在8个任务上超越现有模型,零样本分类准确率达82.7%
- 适合皮肤病学研究、多模态AI开发及临床辅助系统构建者
视觉语言模型的兴起推动了医疗AI的飞跃,但皮肤病学因缺乏标准图像-文本配对而进展缓慢。现有数据集规模与深度不足,仅提供单一标签且缺少临床背景。为此,我们提出Derm1M,首个大规模皮肤病视觉语言数据集,包含1,029,761张图像-文本对。数据源自多样教育资料,围绕专家协作构建的标准本体结构,涵盖4层级、超过390种皮肤状况及130个临床概念,提供病史、症状、肤色等丰富上下文信息。为验证其潜力,我们在该数据集上预训练系列类似CLIP的模型(统称DermLIP),在8个不同数据集的多项任务中显著优于当前最优基础模型,包括零样本皮肤疾病分类、临床与伪影概念识别、少样本/全样本学习以及跨模态检索。代码与数据将公开于https://github.com/SiyuanYan1/Derm1M。
原文摘要 · Abstract (English)
The emergence of vision-language models has transformed medical AI, enabling unprecedented advances in diagnostic capability and clinical applications. However, progress in dermatology has lagged behind other medical domains due to the lack of standard image-text pairs. Existing dermatological datasets are limited in both scale and depth, offering only single-label annotations across a narrow range of diseases instead of rich textual descriptions, and lacking the crucial clinical context needed for real-world applications. To address these limitations, we present Derm1M, the first large-scale vision-language dataset for dermatology, comprising 1,029,761 image-text pairs. Built from diverse educational resources and structured around a standard ontology collaboratively developed by experts, Derm1M provides comprehensive coverage for over 390 skin conditions across four hierarchical levels and 130 clinical concepts with rich contextual information such as medical history, symptoms, and skin tone. To demonstrate Derm1M potential in advancing both AI research and clinical application, we pretrained a series of CLIP-like models, collectively called DermLIP, on this dataset. The DermLIP family significantly outperforms state-of-the-art foundation models on eight diverse datasets across multiple tasks, including zero-shot skin disease classification, clinical and artifacts concept identification, few-shot/full-shot learning, and cross-modal retrieval. Our dataset and code will be publicly available at https://github.com/SiyuanYan1/Derm1M upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。