用高效生成的肺结节图像提升医疗数据不足问题
Evaluating Utility of Memory Efficient Medical Image Generation: A Study on Lung Nodule Segmentation
- 设计内存高效的分块去噪扩散模型生成带分割标注的肺结节CT图
- 纯合成数据训练模型达基准Dice分数,加合成数据显著提升性能
- 适合数据稀缺场景下医疗AI模型训练,尤其肺结节研究
医疗影像数据公开获取困难限制了AI模型发展。本文提出一种内存高效的分块去噪扩散概率模型(DDPM),用于生成带有肺结节分割标注的合成CT图像。该方法在两种场景下评估:仅使用合成数据训练分割模型,以及将合成图像用于扩充真实数据。结果表明,仅用合成数据训练的模型达到与真实数据基准相当的Dice分数;在真实数据中加入合成图像后,分割性能显著提升。生成图像展示了在真实数据有限情况下增强医学影像数据集的巨大潜力。
原文摘要 · Abstract (English)
The scarcity of publicly available medical imaging data limits the development of effective AI models. This work proposes a memory-efficient patch-wise denoising diffusion probabilistic model (DDPM) for generating synthetic medical images, focusing on CT scans with lung nodules. Our approach generates high-utility synthetic images with nodule segmentation while efficiently managing memory constraints, enabling the creation of training datasets. We evaluate the method in two scenarios: training a segmentation model exclusively on synthetic data, and augmenting real-world training data with synthetic images. In the first case, models trained solely on synthetic data achieve Dice scores comparable to those trained on real-world data benchmarks. In the second case, augmenting real-world data with synthetic images significantly improves segmentation performance. The generated images demonstrate their potential to enhance medical image datasets in scenarios with limited real-world data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。