arXiv:2606.00967cs.CV2026-06

通过文本和局部分割提示,灵活生成高分辨率3D CT影像。

MedSyn2: Flexible Control of 3D CT Generation via Text and Semantically-Defined Segmentation Prompts

论文配图:MedSyn2: Flexible Control of 3D CT Generation via Text and Semantically-Defined Segmentation Prompts
图 1 · 摘自论文原文
  • 结合文本描述与部分器官分割,实现对异常位置和形状的精准控制。
  • 在医学图像生成任务中,感知与语义指标提升24%(相对),生成结果更真实。
  • 适合需要可控数据增强或模型训练的医疗影像研究者使用。

面向体素级医学影像生成,现有方法通常依赖放射科报告作为文本提示,或要求完整器官分割标注。前者空间控制弱,后者标注成本高。本文提出一种灵活的多模态框架,支持可选的文本报告和局部分割提示。用户只需标注特定解剖结构或病灶区域,并附以文字说明其语义,即可实现精细空间控制。模型采用改进的扩散变换器架构,联合处理图像与分割标记,并引入门控注意力机制有效处理长篇报告。实验表明,该方法在感知与语义评价上达到当前最优水平(如均值FID相对提升24%),生成高分辨率、解剖一致的CT体积数据,且在数据增强任务中显著提升数据效率。放射科医生评估显示生成图像与真实图像高度一致。

原文摘要 · Abstract (English)

Generative models for volumetric medical images have found many applications in medical imaging, ranging from data augmentation to serving as priors for inverse problems. For these applications, generating high-resolution 3D images with strong controllability is essential but remains highly challenging. Existing approaches typically control generation either through radiology reports used as text prompts or through full image segmentation. While text-based prompting is flexible, it provides limited spatial control over the location, shape, and boundary of abnormalities. In contrast, segmentation-based methods receive precise spatial guidance but are restrictive in requiring full-organ annotations. In this work, we propose a flexible multimodal framework for controllable volumetric image generation that supports input from radiology reports and segmentation prompts (both optional). Our approach allows users to provide segmentation of a specific anatomy or abnormality without requiring full-organ annotations. The semantic meaning of the segmentation mask is specified through an accompanying text description, resulting in a highly flexible and scalable conditioning mechanism. We develop a memory-efficient architecture based on a modified diffusion transformer that jointly processes image and segmentation tokens. The model further incorporates gated attention to effectively attend to long radiology reports. Experiments demonstrate that our method achieves state-of-the-art perceptual and semantic scores (e.g., 24% relative improvement in mean FID), generates high-resolution anatomically consistent CT volumes, and improves data efficiency when used for data augmentation. Radiologists' evaluation further confirms strong alignment between generated and real medical images.

3D生成医学影像扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。