arXiv:2502.20667cs.CVcs.AI2025-02被引 7

用微调扩散模型生成高精度医学图像,提升诊断辅助能力

Advancing AI-Powered Medical Image Synthesis: Insights from MedVQA-GI Challenge Using CLIP, Fine-Tuned Stable Diffusion, and Dream-Booth + LoRA

  • 融合微调Stable Diffusion与DreamBooth+LoRA,实现文本生成动态医学影像
  • Stable Diffusion在多中心数据上FID低至0.064,Inception Score达2.327,质量最优
  • 适合医学AI研发者、临床辅助系统开发者关注,推动真实场景落地

MEDVQA-GI挑战聚焦于将AI驱动的文本到图像生成模型应用于医学诊断,旨在通过合成图像增强诊断能力。现有方法多局限于静态图像分析,缺乏从文本描述生成动态医学影像的能力。本研究提出一种新方法,基于微调生成模型,实现从文本描述生成动态、可扩展、高精度的医学图像。任务分为图像合成(IS)和最优提示生成(OPG)两部分:前者通过自然语言提示生成医学图像,后者生成能产出高质量特定类别图像的提示。研究指出传统方法存在手绘限制、数据集受限、流程静态及通用模型不足等缺陷。评估显示,Stable Diffusion在生成质量与多样性上优于CLIP及DreamBooth+LoRA。其在单中心、多中心及合并数据上的弗雷歇初始距离(FID)分别为0.099、0.064和0.067,平均初始得分(Inception Score)为2.327,表现最佳。该成果推进了AI赋能医学诊断的发展。未来研究将聚焦模型优化、数据增强与临床应用中的伦理问题。

原文摘要 · Abstract (English)

The MEDVQA-GI challenge addresses the integration of AI-driven text-to-image generative models in medical diagnostics, aiming to enhance diagnostic capabilities through synthetic image generation. Existing methods primarily focus on static image analysis and lack the dynamic generation of medical imagery from textual descriptions. This study intends to partially close this gap by introducing a novel approach based on fine-tuned generative models to generate dynamic, scalable, and precise images from textual descriptions. Particularly, our system integrates fine-tuned Stable Diffusion and DreamBooth models, as well as Low-Rank Adaptation (LORA), to generate high-fidelity medical images. The problem is around two sub-tasks namely: image synthesis (IS) and optimal prompt production (OPG). The former creates medical images via verbal prompts, whereas the latter provides prompts that produce high-quality images in specified categories. The study emphasizes the limitations of traditional medical image generation methods, such as hand sketching, constrained datasets, static procedures, and generic models. Our evaluation measures showed that Stable Diffusion surpasses CLIP and DreamBooth + LORA in terms of producing high-quality, diversified images. Specifically, Stable Diffusion had the lowest Fréchet Inception Distance (FID) scores (0.099 for single center, 0.064 for multi-center, and 0.067 for combined), indicating higher image quality. Furthermore, it had the highest average Inception Score (2.327 across all datasets), indicating exceptional diversity and quality. This advances the field of AI-powered medical diagnosis. Future research will concentrate on model refining, dataset augmentation, and ethical considerations for efficiently implementing these advances into clinical practice

医学图像生成扩散模型AI诊断Stable Diffusion

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。