构建首个皮肤病多模态数据集,提升医学视觉语言模型诊断能力
MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks
- 从专业教材提取3类影像与近万对图文,构建高质量皮肤病数据集
- 生成27000+指令跟随型问答样本,规模达现有最大数据集9倍
- 推出皮肤专用视觉语言模型SkinVL,显著优于通用与医疗通用模型
医学视觉语言模型在多个临床领域展现潜力,但针对皮肤病的专用模型仍不成熟,主要受限于现有数据集中文本描述的专业性不足。为此,我们提出MM-Skin,首个涵盖临床、皮肤镜和病理三种成像模态的大规模皮肤病多模态数据集,包含近10,000对高质量图像-文本配对,均源自专业教材。此外,我们还生成了超过27,000个多样化、指令遵循型视觉问答(VQA)样本,数量为当前最大皮肤病VQA数据集的9倍。基于公开数据与MM-Skin,我们开发了专用于皮肤疾病分析的SkinVL模型。在8个数据集上的综合评估显示,SkinVL在视觉问答、监督微调(SFT)及零样本分类任务中表现卓越,显著优于通用及医学通用视觉语言模型。MM-Skin与SkinVL的推出,为临床皮肤病视觉语言助手的发展提供了重要支持。数据集已开源:https://github.com/ZwQ803/MM-Skin
原文摘要 · Abstract (English)
Medical vision-language models (VLMs) have shown promise as clinical assistants across various medical fields. However, specialized dermatology VLM capable of delivering professional and detailed diagnostic analysis remains underdeveloped, primarily due to less specialized text descriptions in current dermatology multimodal datasets. To address this issue, we propose MM-Skin, the first large-scale multimodal dermatology dataset that encompasses 3 imaging modalities, including clinical, dermoscopic, and pathological and nearly 10k high-quality image-text pairs collected from professional textbooks. In addition, we generate over 27k diverse, instruction-following vision question answering (VQA) samples (9 times the size of current largest dermatology VQA dataset). Leveraging public datasets and MM-Skin, we developed SkinVL, a dermatology-specific VLM designed for precise and nuanced skin disease interpretation. Comprehensive benchmark evaluations of SkinVL on VQA, supervised fine-tuning (SFT) and zero-shot classification tasks across 8 datasets, reveal its exceptional performance for skin diseases in comparison to both general and medical VLM models. The introduction of MM-Skin and SkinVL offers a meaningful contribution to advancing the development of clinical dermatology VLM assistants. MM-Skin is available at https://github.com/ZwQ803/MM-Skin
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。