arXiv:2411.13852cs.CVcs.LG2024-11NeurIPS被引 5

提出新方法缓解生成图像污染对在线持续学习的负面影响

Dealing with Synthetic Data Contamination in Online Continual Learning

  • 设计基于熵与真实/合成图像相似性最大化的筛选策略
  • 在高污染数据下显著提升在线持续学习模型性能
  • 适合关注数据污染与模型鲁棒性的计算机视觉研究者

图像生成技术,特别是基于扩散模型的方法,在生成高保真逼真图像方面取得了显著进展。然而,人工智能生成图像的普及可能对机器学习领域带来尚未明确识别的副作用。深度学习在计算机视觉中的成功依赖于互联网上大规模数据集的积累。随着网络中合成数据的大量增加,未来研究人员将难以获取不含人工智能生成内容的‘干净’数据集。已有研究表明,使用被合成图像污染的数据集进行训练会导致性能下降。本文研究了污染数据集对在线持续学习(Online Continual Learning, CL)的影响,实验表明污染数据会阻碍现有在线CL方法的训练。为此,我们提出熵选择与真实-合成相似性最大化(ESRM)方法,以缓解合成图像导致的性能退化。实验结果显示,该方法能显著减轻性能下降,尤其在污染严重时效果更明显。为保证可复现性,代码已开源:https://github.com/maorong-wang/ESRM。

原文摘要 · Abstract (English)

Image generation has shown remarkable results in generating high-fidelity realistic images, in particular with the advancement of diffusion-based models. However, the prevalence of AI-generated images may have side effects for the machine learning community that are not clearly identified. Meanwhile, the success of deep learning in computer vision is driven by the massive dataset collected on the Internet. The extensive quantity of synthetic data being added to the Internet would become an obstacle for future researchers to collect "clean" datasets without AI-generated content. Prior research has shown that using datasets contaminated by synthetic images may result in performance degradation when used for training. In this paper, we investigate the potential impact of contaminated datasets on Online Continual Learning (CL) research. We experimentally show that contaminated datasets might hinder the training of existing online CL methods. Also, we propose Entropy Selection with Real-synthetic similarity Maximization (ESRM), a method to alleviate the performance deterioration caused by synthetic images when training online CL models. Experiments show that our method can significantly alleviate performance deterioration, especially when the contamination is severe. For reproducibility, the source code of our work is available at https://github.com/maorong-wang/ESRM.

持续学习生成数据数据污染扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。