用生成模型合成MRI数据,提升脑肿瘤分割边界精度。
Assessment of Using Synthetic Data in Brain Tumor Segmentation
- 用预训练GAN生成合成MRI,混合真实数据训练U-Net
- 40%真实+60%合成数据时整体性能最佳,边界更清晰
- 肿瘤核心和增强部分仍存在类别不平衡问题
基于MRI的脑肿瘤手动分割因肿瘤异质性、标注数据稀缺及医学图像中的类别不平衡而困难。生成模型产生的合成数据有望通过提升数据多样性缓解这些问题。本研究作为概念验证,探讨将预训练GAN模型生成的合成MRI数据融入U-Net分割网络训练的影响。实验使用BraTS 2020的真实数据、medigan库生成的合成数据,以及不同比例的真实与合成数据混合集。尽管真实数据训练与混合数据训练模型在总体定量指标(Dice系数、IoU、精确率、召回率、准确率)上表现相近,但定性分析显示,40%真实+60%合成数据的混合集在整体肿瘤边界勾画上有所改善。然而,肿瘤核心和增强区域的区域级准确率仍较低,表明类别不平衡问题依然存在。结果支持合成数据作为脑肿瘤分割的数据增强策略的可行性,同时强调未来需开展更大规模实验、确保体素数据一致性,并进一步缓解类别不平衡。
原文摘要 · Abstract (English)
Manual brain tumor segmentation from MRI scans is challenging due to tumor heterogeneity, scarcity of annotated data, and class imbalance in medical imaging datasets. Synthetic data generated by generative models has the potential to mitigate these issues by improving dataset diversity. This study investigates, as a proof of concept, the impact of incorporating synthetic MRI data, generated using a pre-trained GAN model, into training a U-Net segmentation network. Experiments were conducted using real data from the BraTS 2020 dataset, synthetic data generated with the medigan library, and hybrid datasets combining real and synthetic samples in varying proportions. While overall quantitative performance (Dice coefficient, IoU, precision, recall, accuracy) was comparable between real-only and hybrid-trained models, qualitative inspection suggested that hybrid datasets, particularly with 40% real and 60% synthetic data, improved whole tumor boundary delineation. However, region-wise accuracy for the tumor core and the enhancing tumor remained lower, indicating a persistent class imbalance. The findings support the feasibility of synthetic data as an augmentation strategy for brain tumor segmentation, while highlighting the need for larger-scale experiments, volumetric data consistency, and mitigating class imbalance in future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。