训练集重复图像会降低图像分类模型准确率和训练效率。
Impact of Data Duplication on Deep Neural Network-Based Image Classifiers: Robust vs. Standard Models
- 研究重复图像对图像分类DNN的影响,对比标准与鲁棒模型表现。
- 非均匀重复显著降低模型准确率,即使均匀重复也难提升性能。
- 对对抗训练模型影响更明显,适合数据清洗与模型评估参考。
深度神经网络图像分类器的准确率和鲁棒性受训练数据质量、模型结构、训练过程及部署环境等因素显著影响。近年来,语言模型中训练数据重复问题引发广泛关注,去重可提升训练性能与模型准确率。尽管图像分类DNN的数据质量重要性已被广泛认可,但训练集中重复图像对模型泛化与性能的影响仍缺乏研究。本文系统分析了图像分类中重复数据的影响,发现训练集中的重复图像不仅降低训练效率,还可能导致分类器准确率下降。该负面影响在类别间重复分布不均时尤为显著,且在对抗训练模型中,无论重复是否均匀,均会加剧性能损失。即使重复样本均匀选取,增加重复量也无法显著提升准确率。
原文摘要 · Abstract (English)
The accuracy and robustness of machine learning models against adversarial attacks are significantly influenced by factors such as training data quality, model architecture, the training process, and the deployment environment. In recent years, duplicated data in training sets, especially in language models, has attracted considerable attention. It has been shown that deduplication enhances both training performance and model accuracy in language models. While the importance of data quality in training image classifier Deep Neural Networks (DNNs) is widely recognized, the impact of duplicated images in the training set on model generalization and performance has received little attention. In this paper, we address this gap and provide a comprehensive study on the effect of duplicates in image classification. Our analysis indicates that the presence of duplicated images in the training set not only negatively affects the efficiency of model training but also may result in lower accuracy of the image classifier. This negative impact of duplication on accuracy is particularly evident when duplicated data is non-uniform across classes or when duplication, whether uniform or non-uniform, occurs in the training set of an adversarially trained model. Even when duplicated samples are selected in a uniform way, increasing the amount of duplication does not lead to a significant improvement in accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。