arXiv:2509.05978eess.IVcs.CL2025-09中稿 · the 2025 MICCAI EL…被引 1

用自然语言生成高分辨率3D医学影像,模拟疾病变化过程。

Imagining Alternatives: Towards High-Resolution 3D Counterfactual Medical Image Generation via Language Guidance

  • 基于3D扩散模型,结合语言提示生成合成患者影像。
  • 在多发性硬化和阿尔茨海默病数据集上成功模拟病变负荷与认知状态变化。
  • 首个实现原生3D语言引导医学影像生成,适合临床研究与教学使用。

视觉-语言模型在2D图像生成中表现优异,但其成功依赖于大量现成的预训练基础模型。然而,3D领域尚无类似预训练模型,严重制约了发展。因此,仅通过自然语言条件生成高分辨率3D反事实医学图像的潜力尚未被探索。本工作提出一个框架,可基于自由文本提示生成高分辨率3D反事实医学图像,用于模拟合成患者的影像。我们改进了先进的3D扩散模型,融合Simple Diffusion优化并增强条件输入以提升文本对齐与图像质量。据我们所知,这是首个将语言引导的原生3D扩散模型应用于神经影像的研究,其中三维建模至关重要。在两个神经影像MRI数据集上,该框架成功模拟了多发性硬化患者的病变负荷变化及阿尔茨海默病患者的认知状态差异,生成高质量图像的同时保持个体真实性。结果为3D医学影像中的提示驱动疾病进展分析奠定基础。

原文摘要 · Abstract (English)

Vision-language models have demonstrated impressive capabilities in generating 2D images under various conditions; however, the success of these models is largely enabled by extensive, readily available pretrained foundation models. Critically, comparable pretrained models do not exist for 3D, significantly limiting progress. As a result, the potential of vision-language models to produce high-resolution 3D counterfactual medical images conditioned solely on natural language remains unexplored. Addressing this gap would enable powerful clinical and research applications, such as personalized counterfactual explanations, simulation of disease progression, and enhanced medical training by visualizing hypothetical conditions in realistic detail. Our work takes a step toward this challenge by introducing a framework capable of generating high-resolution 3D counterfactual medical images of synthesized patients guided by free-form language prompts. We adapt state-of-the-art 3D diffusion models with enhancements from Simple Diffusion and incorporate augmented conditioning to improve text alignment and image quality. To our knowledge, this is the first demonstration of a language-guided native-3D diffusion model applied to neurological imaging, where faithful three-dimensional modeling is essential. On two neurological MRI datasets, our framework simulates varying counterfactual lesion loads in Multiple Sclerosis and cognitive states in Alzheimer's disease, generating high-quality images while preserving subject fidelity. Our results lay the groundwork for prompt-driven disease progression analysis in 3D medical imaging. Project link - https://lesupermomo.github.io/imagining-alternatives/.

3D生成医学影像语言引导反事实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。