用文本提示生成配对的医学图像与分割图,解决数据稀缺问题。
MedSegFactory: Text-Guided Generation of Medical Image-Mask Pairs
- 双流扩散模型协同生成图像与掩码,通过动态交叉注意力对齐。
- 支持用户自定义标签、模态和病灶条件,按需生成高质量配对数据。
- 可直接用于训练分割模型,适合医疗影像数据增强与算法开发。
本文提出MedSegFactory,一个跨模态、跨任务的医学图像-掩码成对合成框架,旨在构建无限的数据资源库,为现有分割工具提供图像-掩码对。其核心是双流扩散模型:一流生成医学图像,另一流生成对应分割掩码。为确保配对精确,引入联合交叉注意力(JCA),实现两流间的动态交叉条件化,形成双向引导的去噪机制,使图像与掩码相互指导生成,提升一致性。用户可通过指定目标标签、成像模态、解剖区域和病理状态等文本提示,实现按需生成。实验表明,MedSegFactory在2D与3D分割任务中均达到竞争性或领先性能,有效缓解数据稀缺与监管限制,推动医学影像工作流程高效化与精准化。
原文摘要 · Abstract (English)
This paper presents MedSegFactory, a versatile medical synthesis framework that generates high-quality paired medical images and segmentation masks across modalities and tasks. It aims to serve as an unlimited data repository, supplying image-mask pairs to enhance existing segmentation tools. The core of MedSegFactory is a dual-stream diffusion model, where one stream synthesizes medical images and the other generates corresponding segmentation masks. To ensure precise alignment between image-mask pairs, we introduce Joint Cross-Attention (JCA), enabling a collaborative denoising paradigm by dynamic cross-conditioning between streams. This bidirectional interaction allows both representations to guide each other's generation, enhancing consistency between generated pairs. MedSegFactory unlocks on-demand generation of paired medical images and segmentation masks through user-defined prompts that specify the target labels, imaging modalities, anatomical regions, and pathological conditions, facilitating scalable and high-quality data generation. This new paradigm of medical image synthesis enables seamless integration into diverse medical imaging workflows, enhancing both efficiency and accuracy. Extensive experiments show that MedSegFactory generates data of superior quality and usability, achieving competitive or state-of-the-art performance in 2D and 3D segmentation tasks while addressing data scarcity and regulatory constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。