用大模型生成医学图像细节描述,提升生成真实度与多样性
VAP-Diffusion: Enriching Descriptions with MLLMs for Enhanced Medical Image Generation
- 用思维链提示词从多模态大模型获取精准医学描述
- 在四种数据集上生成图像质量显著优于基线方法
- 适合需要高质量医学图像生成的研究者使用
由于医学图像的外观受多种潜在因素影响,生成模型需超越标签的丰富属性信息才能产出真实且多样化的图像。例如,生成特定形态的皮肤病变图像需包含形状、大小、纹理和颜色等详细描述,但这些信息常不可得。为此,我们提出视觉属性提示框架(VAP-Diffusion),利用预训练多模态大模型(MLLMs)提取外部知识以增强医学图像生成质量与多样性。首先,针对皮肤病、结直肠及胸部X光等常见影像任务,设计遵循思维链(Chain-of-Thought)的提示策略,生成无幻觉的描述,并在不同类别中存储。推理时,从对应类别随机检索描述用于生成。此外,为应对测试时未见描述组合,提出原型条件机制,限制测试嵌入与训练嵌入相似。在四个数据集上的三类医学影像实验验证了该方法的有效性。
原文摘要 · Abstract (English)
As the appearance of medical images is influenced by multiple underlying factors, generative models require rich attribute information beyond labels to produce realistic and diverse images. For instance, generating an image of skin lesion with specific patterns demands descriptions that go beyond diagnosis, such as shape, size, texture, and color. However, such detailed descriptions are not always accessible. To address this, we explore a framework, termed Visual Attribute Prompts (VAP)-Diffusion, to leverage external knowledge from pre-trained Multi-modal Large Language Models (MLLMs) to improve the quality and diversity of medical image generation. First, to derive descriptions from MLLMs without hallucination, we design a series of prompts following Chain-of-Thoughts for common medical imaging tasks, including dermatologic, colorectal, and chest X-ray images. Generated descriptions are utilized during training and stored across different categories. During testing, descriptions are randomly retrieved from the corresponding category for inference. Moreover, to make the generator robust to unseen combination of descriptions at the test time, we propose a Prototype Condition Mechanism that restricts test embeddings to be similar to those from training. Experiments on three common types of medical imaging across four datasets verify the effectiveness of VAP-Diffusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。