用文本控制生成皮肤病图像和分割图,提升诊断模型性能。
SkinDualGen: Prompt-Driven Diffusion for Simultaneous Image-Mask Generation in Skin Lesions
- 基于Stable Diffusion-2.0,通过LoRA微调实现图文条件生成。
- 合成数据使分类与分割模型准确率提升8%~15%,F1-score显著提高。
- 适合医学图像数据稀缺场景,尤其适用于罕见病诊断研究。
医学影像分析在皮肤病变等疾病的早期诊断中至关重要,但数据稀缺和类别不平衡严重制约深度学习模型表现。本文提出一种新方法,利用预训练的Stable Diffusion-2.0模型生成高质量合成皮肤病变图像及其对应分割掩码,用于增强分类与分割任务的训练数据。通过领域特定的低秩适配(LoRA)微调及多目标损失函数联合优化,模型可单步生成符合文本描述的临床相关图像与掩码。实验表明,生成图像经FID评分验证,质量接近真实图像。混合真实与合成数据集显著提升分类与分割模型性能,准确率与F1-score提升8%至15%,其他关键指标如Dice系数和IoU亦获正向改善。该方法为医学影像数据挑战提供可扩展解决方案,有助于提升罕见病诊断的准确性与可靠性。
原文摘要 · Abstract (English)
Medical image analysis plays a pivotal role in the early diagnosis of diseases such as skin lesions. However, the scarcity of data and the class imbalance significantly hinder the performance of deep learning models. We propose a novel method that leverages the pretrained Stable Diffusion-2.0 model to generate high-quality synthetic skin lesion images and corresponding segmentation masks. This approach augments training datasets for classification and segmentation tasks. We adapt Stable Diffusion-2.0 through domain-specific Low-Rank Adaptation (LoRA) fine-tuning and joint optimization of multi-objective loss functions, enabling the model to simultaneously generate clinically relevant images and segmentation masks conditioned on textual descriptions in a single step. Experimental results show that the generated images, validated by FID scores, closely resemble real images in quality. A hybrid dataset combining real and synthetic data markedly enhances the performance of classification and segmentation models, achieving substantial improvements in accuracy and F1-score of 8% to 15%, with additional positive gains in other key metrics such as the Dice coefficient and IoU. Our approach offers a scalable solution to address the challenges of medical imaging data, contributing to improved accuracy and reliability in diagnosing rare diseases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。