用文本生成结肠镜图像,解决医疗数据少难题。
Prompt to Polyp: Medical Text-Conditioned Image Synthesis with Diffusion Models
- 用临床文本编码器+交叉注意力优化扩散模型,对齐医学描述与图像。
- 自研小模型MSDM在两项医疗数据集上达到接近大模型的生成质量。
- 适合医疗数据稀缺场景,兼顾效果与计算成本,医生可参与评估。
从文本描述生成真实医疗图像,在缓解医疗AI数据稀缺问题的同时,有助于保护患者隐私。本文系统研究了医学领域的文本到图像合成,对比两种方法:(1) 微调大型预训练潜在扩散模型,(2) 训练小型专用模型。提出一种新模型MSDM,基于Stable Diffusion架构,集成临床文本编码器、变分自编码器和交叉注意力机制,更好地对齐医学文本提示与生成图像。在结肠镜(MedVQA-GI)和放射科(ROCOv2)数据集上的评估显示,尽管大型模型在保真度上更优,但优化后的MSDM在生成质量上可媲美,且计算成本更低。定量指标与医学专家的定性评估揭示了两种方法的优势与局限。
原文摘要 · Abstract (English)
The generation of realistic medical images from text descriptions has significant potential to address data scarcity challenges in healthcare AI while preserving patient privacy. This paper presents a comprehensive study of text-to-image synthesis in the medical domain, comparing two distinct approaches: (1) fine-tuning large pre-trained latent diffusion models and (2) training small, domain-specific models. We introduce a novel model named MSDM, an optimized architecture based on Stable Diffusion that integrates a clinical text encoder, variational autoencoder, and cross-attention mechanisms to better align medical text prompts with generated images. Our study compares two approaches: fine-tuning large pre-trained models (FLUX, Kandinsky) versus training compact domain-specific models (MSDM). Evaluation across colonoscopy (MedVQA-GI) and radiology (ROCOv2) datasets reveals that while large models achieve higher fidelity, our optimized MSDM delivers comparable quality with lower computational costs. Quantitative metrics and qualitative evaluations by medical experts reveal strengths and limitations of each approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。