用生成模型合成稀缺MRI影像数据,提升罕见模态分割精度
MRGen: Segmentation Data Engine for Underrepresented MRI Modalities

- 基于扩散模型,输入文本和掩码生成可控的MRI图像
- 在无标注模态上提升分割性能,最高达28.6%的Dice增益
- 适合医疗数据少、标注难的医学影像研究者使用
针对罕见但临床重要的医学影像模态因标注数据稀缺导致分割模型训练困难的问题,本文提出利用生成模型合成数据以支持分割训练。贡献有三:(i) 构建大规模放射科图像-文本数据集MRGen-DB,包含丰富元数据(模态标签、属性、区域、器官信息),部分样本附带像素级掩码;(ii) 提出基于扩散模型的MRGen数据引擎,可依据文本提示和分割掩码生成多样化且真实的无标注MRI模态图像,适用于低资源场景下的分割训练;(iii) 在多个模态上的大量实验表明,借助高质量合成数据,MRGen显著提升未标注模态的分割性能,最高实现28.6%的Dice系数提升。该方法有效填补了医学图像分析中数据匮乏的空白,拓展了难以获取人工标注场景下的分割能力。代码、模型与数据将公开于https://haoningwu3639.github.io/MRGen/
原文摘要 · Abstract (English)
Training medical image segmentation models for rare yet clinically important imaging modalities is challenging due to the scarcity of annotated data, and manual mask annotations can be costly and labor-intensive to acquire. This paper investigates leveraging generative models to synthesize data, for training segmentation models for underrepresented modalities, particularly on annotation-scarce MRI. Concretely, our contributions are threefold: (i) we introduce MRGen-DB, a large-scale radiology image-text dataset comprising extensive samples with rich metadata, including modality labels, attributes, regions, and organs information, with a subset featuring pixel-wise mask annotations; (ii) we present MRGen, a diffusion-based data engine for controllable medical image synthesis, conditioned on text prompts and segmentation masks. MRGen can generate realistic images for diverse MRI modalities lacking mask annotations, facilitating segmentation training in low-source domains; (iii) extensive experiments across multiple modalities demonstrate that MRGen significantly improves segmentation performance on unannotated modalities by providing high-quality synthetic data. We believe that our method bridges a critical gap in medical image analysis, extending segmentation capabilities to scenarios that are challenging to acquire manual annotations. The codes, models, and data will be publicly available at https://haoningwu3639.github.io/MRGen/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。