研究自生成数据如何在多代训练中放大图像分类模型的偏见
Will the Inclusion of Generated Data Amplify Bias Across Generations in Future Image Classification Models?
- 构建自循环训练环境,模拟生成模型与分类模型协同进化
- 在多个数据集上发现跨代次公平性指标持续下降
- 揭示合成数据引发偏见累积的机制,适合关注AI公平性的研究者
随着对高质量训练数据需求的增长,研究者越来越多地使用生成模型创建合成数据以解决数据稀缺问题,并支持模型持续改进。然而,依赖自生成数据带来一个关键问题:该做法是否会加剧未来模型中的偏见?尽管多数研究聚焦于整体性能,但对模型偏见(尤其是子群体偏见)的影响仍缺乏深入探讨。本文研究了生成数据对图像分类任务的影响,重点关注偏见问题。我们构建了一个实用的模拟环境,包含自消耗循环,使生成模型与分类模型协同训练。在Colorized MNIST、CIFAR-20/100和Hard ImageNet数据集上进行了数百次实验,揭示了跨代次公平性度量的变化。此外,我们提出一个猜想,解释在连续增强数据集上训练模型时偏见动态演化的原因。研究结果为合成数据在真实应用中对公平性影响的持续争论提供了重要依据。
原文摘要 · Abstract (English)
As the demand for high-quality training data escalates, researchers have increasingly turned to generative models to create synthetic data, addressing data scarcity and enabling continuous model improvement. However, reliance on self-generated data introduces a critical question: Will this practice amplify bias in future models? While most research has focused on overall performance, the impact on model bias, particularly subgroup bias, remains underexplored. In this work, we investigate the effects of the generated data on image classification tasks, with a specific focus on bias. We develop a practical simulation environment that integrates a self-consuming loop, where the generative model and classification model are trained synergistically. Hundreds of experiments are conducted on Colorized MNIST, CIFAR-20/100, and Hard ImageNet datasets to reveal changes in fairness metrics across generations. In addition, we provide a conjecture to explain the bias dynamics when training models on continuously augmented datasets across generations. Our findings contribute to the ongoing debate on the implications of synthetic data for fairness in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。