用生成模型合成医学图像,尝试扩充小样本数据集
Using Synthetic Images to Augment Small Medical Image Datasets
- 改进StyleGAN2生成多模态高分辨率医学图像
- 合成图像未提升语义分割模型性能
- 适合研究数据增强在医疗影像中的有效性
近年来,深度学习在医学影像领域受到广泛关注。然而,其性能依赖大规模标注数据集,而多数医学影像数据集规模小,标注样本有限,主要因肿瘤学家勾画图像耗时费力。现有数据增强技术包括仿射变换、弹性变形或通过生成对抗网络(GAN)生成合成图像。本文提出一种新型条件式StyleGAN2方法,用于生成多模态高分辨率医学图像,以扩充六个数据集的小样本训练集。利用真实与合成图像训练语义分割模型,并评估生成图像质量及对分割性能的影响。结果表明,下游分割模型并未因合成图像的加入而获得性能提升,后续需进一步分析此类增强对分割效果的影响。
原文摘要 · Abstract (English)
Recent years have witnessed a growing academic and industrial interest in deep learning (DL) for medical imaging. To perform well, DL models require very large labeled datasets. However, most medical imaging datasets are small, with a limited number of annotated samples. The reason they are small is usually because delineating medical images is time-consuming and demanding for oncologists. There are various techniques that can be used to augment a dataset, for example, to apply affine transformations or elastic transformations to available images, or to add synthetic images generated by a Generative Adversarial Network (GAN). In this work, we have developed a novel conditional variant of a current GAN method, the StyleGAN2, to generate multi-modal high-resolution medical images with the purpose to augment small medical imaging datasets with these synthetic images. We use the synthetic and real images from six datasets to train models for the downstream task of semantic segmentation. The quality of the generated medical images and the effect of this augmentation on the segmentation performance were evaluated afterward. Finally, the results indicate that the downstream segmentation models did not benefit from the generated images. Further work and analyses are required to establish how this augmentation affects the segmentation performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。