研究医学图像复杂度对生成模型性能的影响,揭示数据量与图像质量的关系。
Medical Imaging Complexity and its Effects on GAN Performance
- 用信息论中的delentropy衡量图像复杂度,分析其对生成效果的影响。
- 数据量越大,生成图像越逼真;但图像越复杂,模型性能越差。
- 适用于需要高质量医学图像合成的研究者和医疗AI开发者。
机器学习在临床应用中的普及带来了对高保真医学图像训练数据的迫切需求,但受限于成本和隐私问题,真实数据往往稀缺。基于生成对抗网络(GAN)的医学图像合成技术因此成为一种有效手段,可从现有真实医学图像中生成逼真图像。然而,高效训练GAN所需的数据集规模尚不明确。本文通过实验建立基准,量化了样本数据集大小与生成图像保真度之间的关系,并考察了数据分布中图像复杂度的影响。我们采用源自香农信息论的delentropy作为图像复杂度指标,对两种前沿GAN模型——StyleGAN 3与SPADE-GAN——在多个具有不同样本规模的医学影像数据集上进行训练。结果表明,随着训练集规模增加,生成性能普遍提升;但图像复杂度越高,模型表现越差。
原文摘要 · Abstract (English)
The proliferation of machine learning models in diverse clinical applications has led to a growing need for high-fidelity, medical image training data. Such data is often scarce due to cost constraints and privacy concerns. Alleviating this burden, medical image synthesis via generative adversarial networks (GANs) emerged as a powerful method for synthetically generating photo-realistic images based on existing sets of real medical images. However, the exact image set size required to efficiently train such a GAN is unclear. In this work, we experimentally establish benchmarks that measure the relationship between a sample dataset size and the fidelity of the generated images, given the dataset's distribution of image complexities. We analyze statistical metrics based on delentropy, an image complexity measure rooted in Shannon's entropy in information theory. For our pipeline, we conduct experiments with two state-of-the-art GANs, StyleGAN 3 and SPADE-GAN, trained on multiple medical imaging datasets with variable sample sizes. Across both GANs, general performance improved with increasing training set size but suffered with increasing complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。