arXiv:2503.23623cs.CV2025-03CVPR被引 1

用语言控制医学图像生成,实现关键特征的精准分离与调控

Language-Guided Trajectory Traversal in Disentangled Stable Diffusion Latent Space for Factorized Medical Image Generation

  • 通过语言引导的潜空间轨迹遍历,分离医学图像关键属性
  • 在胸片和皮肤数据集上验证了模型对解耦特征的有效生成能力
  • 适合医疗影像生成、可解释性研究与疾病机制探索的研究者

文本到图像的扩散模型已展现出从自然语言提示生成高分辨率逼真图像的强大能力,这些语言引导的合成图像对疾病可解释性或因果关系探索至关重要。然而,其在医学影像等专业领域中解耦并控制潜在变化因素的潜力仍待挖掘。本文首次探究微调后的视觉-语言基础模型在医学图像数据集上进行潜空间解耦的能力,实现因子化医学图像生成与插值。在胸片和皮肤数据集上的大量实验表明,微调后的语言引导Stable Diffusion能天然学习到图像生成的关键属性解耦,如患者解剖结构或疾病诊断特征。我们提出一种框架,通过潜空间轨迹遍历识别、隔离并操控关键属性,从而实现对医学图像合成的精确控制。

原文摘要 · Abstract (English)

Text-to-image diffusion models have demonstrated a remarkable ability to generate photorealistic images from natural language prompts. These high-resolution, language-guided synthesized images are essential for the explainability of disease or exploring causal relationships. However, their potential for disentangling and controlling latent factors of variation in specialized domains like medical imaging remains under-explored. In this work, we present the first investigation of the power of pre-trained vision-language foundation models, once fine-tuned on medical image datasets, to perform latent disentanglement for factorized medical image generation and interpolation. Through extensive experiments on chest X-ray and skin datasets, we illustrate that fine-tuned, language-guided Stable Diffusion inherently learns to factorize key attributes for image generation, such as the patient's anatomical structures or disease diagnostic features. We devise a framework to identify, isolate, and manipulate key attributes through latent space trajectory traversal of generative models, facilitating precise control over medical image synthesis.

医学图像扩散模型潜空间语言控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。