用生成模型合成图像,能替代真实数据提升分类效果吗?
Exploring the Equivalence of Closed-Set Generative and Real Data Augmentation in Image Classification
- 用生成模型在给定数据集上合成图像,用于增强分类训练
- 实测得出:合成图像需约3倍规模才能达到真实数据增广效果
- 适用于数据稀缺场景,尤其对医学图像等小样本任务有指导意义
本文探讨机器学习中的关键问题:给定一个图像分类任务的训练集,能否通过在此数据集上训练生成模型来提升分类性能?(即闭集生成数据增广)。我们首先分析真实图像与先进生成模型所产闭集合成图像之间的异同。通过大量实验,系统揭示了闭集合成数据用于增广的有效性。特别地,我们实证确定了实现等效增广所需的合成图像规模。此外,还量化证明了真实数据增广与开集生成增广(使用外部数据训练的生成模型)之间存在等价关系。尽管直观上真实图像更优,但我们的实证结果提供了一种量化指南,说明需扩大合成数据规模以达到相似分类表现。在自然图像与医学图像数据集上的结果进一步表明,该效应随基础训练集大小和合成数据量变化而动态调整。
原文摘要 · Abstract (English)
In this paper, we address a key scientific problem in machine learning: Given a training set for an image classification task, can we train a generative model on this dataset to enhance the classification performance? (i.e., closed-set generative data augmentation). We start by exploring the distinctions and similarities between real images and closed-set synthetic images generated by advanced generative models. Through extensive experiments, we offer systematic insights into the effective use of closed-set synthetic data for augmentation. Notably, we empirically determine the equivalent scale of synthetic images needed for augmentation. In addition, we also show quantitative equivalence between the real data augmentation and open-set generative augmentation (generative models trained using data beyond the given training set). While it aligns with the common intuition that real images are generally preferred, our empirical formulation also offers a guideline to quantify the increased scale of synthetic data augmentation required to achieve comparable image classification performance. Our results on natural and medical image datasets further illustrate how this effect varies with the baseline training set size and the amount of synthetic data incorporated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。