用生成模型扩充脑瘤MRI数据,效果因模型和比例而异。
Do Synthetic Brain MRIs Reliably Improve Tumour Classification? A StyleGAN2-ADA Class-Plane Augmentation Study on BRISC 2025

- 用12个类平面StyleGAN2-ADA生成假影像并加入真实数据
- MobileViTV2在1:1比例下准确率提升1.02%,统计显著
- 生成图像质量高但未必有用,效果取决于具体模型和比例
生成式增强常被用于小样本医疗图像数据集,但只有当它能提升下游任务性能时才真正有用。本文采用生成补充方式:将GAN生成的样本添加至真实训练集,而非对现有图像进行几何或光度变换。在受限的BRISC 2025数据子集上训练了12个类平面StyleGAN2-ADA生成器,测试其输出(有无InceptionV3特征空间过滤)在三种分类器上的表现:基于InceptionV3特征的随机森林(RF)、紧凑型双头卷积神经网络(CNN)以及MobileViTV2(轻量级混合卷积-注意力模型)。评估在1:1和1:2真实与合成图像比例下进行。独立的GPT-5.5盲测显示,模型可辨识真实与合成图像的准确率为57.73%(95%置信区间:54.48–60.92%),仅略高于随机水平。随机森林未从合成数据中受益;CNN虽有平均提升,但未通过Holm校正;而MobileViTV2表现最明显:经过滤的1:1增广使肿瘤分类准确率绝对提升1.02%(95%置信区间:0.54–1.54%;校正后p=0.0104)。另一次效率分析发现,所有增广的CNN条件均比基线提前42–64%选择最佳检查点,而计算量匹配的MobileViTV2运行也提前50–67%完成真实数据训练。总体而言,增广效果依赖于模型架构与合成比例,并非仅由视觉保真度决定。
原文摘要 · Abstract (English)
Generative augmentation is often proposed as a remedy for small medical-image datasets, but synthetic images are only useful when they improve downstream task performance. "Augmentation" here means synthetic supplementation: GAN-generated samples added to the real training pool, not geometric or photometric transforms of existing images. Twelve class-plane StyleGAN2-ADA generators were trained on constrained BRISC 2025 partitions to test whether their output, with or without InceptionV3 feature-space filtering, improves held-out tumour classification across three classifier families: a random forest (RF) on InceptionV3 features, a compact two-headed convolutional neural network (CNN), and MobileViTV2, a mobile hybrid convolutional-transformer. Each was evaluated at 1:1 and 1:2 real-to-synthetic ratios. An independent GPT-5.5 blind test placed gated real-versus-synthetic discrimination at 57.73% (95% CI: 54.48--60.92%) on the model-legible subset -- modestly above chance. The RF classifier did not benefit from the synthetic MRIs. The CNN showed consistent mean gains that did not survive Holm correction. MobileViTV2 showed the clearest benefit: filtered 1:1 augmentation improved tumour classification accuracy by 1.02% absolute (95% CI: 0.54--1.54%; Holm-corrected p = 0.0104). A secondary efficiency analysis found that every augmented CNN condition selected its checkpoint 42--64% earlier than baseline, while compute-matched MobileViTV2 runs reached selection after 50--67% fewer real-data epochs. Overall, augmentation utility was found to be architecture- and ratio-dependent, not guaranteed by visual fidelity alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。