arXiv:2506.20449cs.CV2025-06被引 4

用文本生成医学图像,小数据也能出好图。

Med-Art: Diffusion Transformer for 2D Medical Text-to-Image Generation

  • 基于DiT架构,用视觉语言模型补足医疗文本数据不足
  • 在两个数据集上FID、KID指标领先,分类任务表现优
  • 提出混合层级微调方法,解决色彩过饱和问题

近年来,文本到图像生成模型取得了显著进展。然而,在医学图像生成领域仍面临数据集规模小、医疗文本数据稀缺等挑战。为此,我们提出Med-Art框架,专为小样本医疗图像生成设计。Med-Art利用视觉语言模型生成医学图像的视觉描述,缓解医疗文本数据不足的问题。该框架基于大规模预训练文本到图像模型PixArt-α(采用扩散Transformer,DiT),在有限数据下实现高性能。此外,我们提出创新的混合层级扩散微调(HLDF)方法,引入像素级损失,有效改善图像过饱和等问题。在两个医学图像数据集上,我们的方法在FID、KID及下游分类任务中均达到当前最优性能。

原文摘要 · Abstract (English)

Text-to-image generative models have achieved remarkable breakthroughs in recent years. However, their application in medical image generation still faces significant challenges, including small dataset sizes, and scarcity of medical textual data. To address these challenges, we propose Med-Art, a framework specifically designed for medical image generation with limited data. Med-Art leverages vision-language models to generate visual descriptions of medical images which overcomes the scarcity of applicable medical textual data. Med-Art adapts a large-scale pre-trained text-to-image model, PixArt-$α$, based on the Diffusion Transformer (DiT), achieving high performance under limited data. Furthermore, we propose an innovative Hybrid-Level Diffusion Fine-tuning (HLDF) method, which enables pixel-level losses, effectively addressing issues such as overly saturated colors. We achieve state-of-the-art performance on two medical image datasets, measured by FID, KID, and downstream classification performance.

文本生成图像医学图像扩散模型小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。