用扩散模型生成医学影像,让CNN在合成数据上表现接近真实数据。
Diffusion-Based Approaches in Medical Image Generation and Analysis
- 用扩散模型生成脑肿瘤MRI、白血病和新冠CT的合成图像。
- 三种数据集的CNN分类性能均达到与真实数据训练相当水平。
- 适合想减少真实患者数据依赖的研究者或医疗AI开发者。
医学影像数据因隐私问题存在稀缺性,扩散模型作为新兴生成技术,可生成逼真合成数据以缓解此问题。本研究考察了扩散模型在三个领域生成合成医学图像的有效性:脑肿瘤MRI、急性淋巴细胞白血病(ALL)和SARS-CoV-2 CT扫描。针对每个领域,训练了一个扩散模型生成合成数据集,并使用预训练的CNN架构在这些合成数据上进行训练,随后在未见的真实数据上评估性能。结果表明,三种数据集上的分类任务均取得了令人满意的表现。通过局部可解释模型无关解释(LIME)分析发现,模型关注的是与分类相关的图像特征。该研究证明了扩散模型在生成可用于训练CNN的医学图像方面的潜力。
原文摘要 · Abstract (English)
Data scarcity in medical imaging poses significant challenges due to privacy concerns. Diffusion models, a recent generative modeling technique, offer a potential solution by generating synthetic and realistic data. However, questions remain about the performance of convolutional neural network (CNN) models on original and synthetic datasets. If diffusion-generated samples can help CNN models perform comparably to those trained on original datasets, reliance on patient-specific data for training CNNs might be reduced. In this study, we investigated the effectiveness of diffusion models for generating synthetic medical images to train CNNs in three domains: Brain Tumor MRI, Acute Lymphoblastic Leukemia (ALL), and SARS-CoV-2 CT scans. A diffusion model was trained to generate synthetic datasets for each domain. Pre-trained CNN architectures were then trained on these synthetic datasets and evaluated on unseen real data. All three datasets achieved promising classification performance using CNNs trained on synthetic data. Local Interpretable Model-Agnostic Explanations (LIME) analysis revealed that the models focused on relevant image features for classification. This study demonstrates the potential of diffusion models to generate synthetic medical images for training CNNs in medical image analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。