用开源皮肤数据生成9万多对图文,助力皮肤病AI研究
DermaSynth: Rich Synthetic Image-Text Pairs Using Open Access Dermatology Datasets
- 用Gemini 2.0和自指导方法生成多样文本,结合元数据降幻觉
- 构建含92,020对图文的公开数据集,覆盖13,568临床与35,561皮肤镜图像
- 适配皮肤病视觉大模型训练,支持科研与临床应用
开发皮肤病视觉大语言模型的一大障碍是缺乏大规模图像-文本配对数据集。本文提出DermaSynth,一个包含92,020个合成图像-文本对的数据集,基于45,205张图像(13,568张临床图像和35,561张皮肤镜图像)构建,用于皮肤病相关临床任务。利用先进的LLM Gemini 2.0,通过临床相关提示词与自指导方法生成丰富多样的合成文本,并在输入提示中融入数据集元信息以减少潜在幻觉。该数据集基于具有宽松CC-BY-4.0许可的开源皮肤图像库(DERM12345、BCN20000、PAD-UFES-20、SCIN 和 HIBA)。我们还基于5,000个样本微调了初步的Llama-3.2-11B-Vision-Instruct模型,命名为DermatoLlama 1.0。本工作成果可促进并加速皮肤病领域的AI研究。数据与代码已公开于https://github.com/abdurrahimyilmaz/DermaSynth。
原文摘要 · Abstract (English)
A major barrier to developing vision large language models (LLMs) in dermatology is the lack of large image--text pairs dataset. We introduce DermaSynth, a dataset comprising of 92,020 synthetic image--text pairs curated from 45,205 images (13,568 clinical and 35,561 dermatoscopic) for dermatology-related clinical tasks. Leveraging state-of-the-art LLMs, using Gemini 2.0, we used clinically related prompts and self-instruct method to generate diverse and rich synthetic texts. Metadata of the datasets were incorporated into the input prompts by targeting to reduce potential hallucinations. The resulting dataset builds upon open access dermatological image repositories (DERM12345, BCN20000, PAD-UFES-20, SCIN, and HIBA) that have permissive CC-BY-4.0 licenses. We also fine-tuned a preliminary Llama-3.2-11B-Vision-Instruct model, DermatoLlama 1.0, on 5,000 samples. We anticipate this dataset to support and accelerate AI research in dermatology. Data and code underlying this work are accessible at https://github.com/abdurrahimyilmaz/DermaSynth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。