arXiv:2412.02366cs.CV2024-12被引 16

用生成模型编辑图像,提升跨域分类效果

GenMix: Effective Data Augmentation with Generative Diffusion Model Image Editing

  • 根据提示词编辑图像,生成混合新样本
  • 在8个数据集上均提升分类准确率,跨域效果更优
  • 适合数据少或需对抗鲁棒性的场景

数据增强广泛用于提升视觉分类任务的泛化能力。然而,传统方法在源域与目标域差异较大时表现不佳,难以弥合领域差距。本文提出GenMix,一种通用的提示引导生成式数据增强方法,可提升域内及跨域图像分类性能。该方法基于自定义条件提示对图像进行编辑,生成增强图像,并通过融合输入图像与生成图像的局部区域,结合分形图案,减少不真实图像和标签模糊问题,从而提升模型性能与对抗鲁棒性。在八个公开数据集上的大量实验验证了其有效性,涵盖通用与细粒度分类,在域内与跨域设置下均有提升。此外,该方法在自监督学习、数据稀缺场景及对抗鲁棒性方面也表现出色。相比现有最先进方法,本方法在各项任务中均表现更优。

原文摘要 · Abstract (English)

Data augmentation is widely used to enhance generalization in visual classification tasks. However, traditional methods struggle when source and target domains differ, as in domain adaptation, due to their inability to address domain gaps. This paper introduces GenMix, a generalizable prompt-guided generative data augmentation approach that enhances both in-domain and cross-domain image classification. Our technique leverages image editing to generate augmented images based on custom conditional prompts, designed specifically for each problem type. By blending portions of the input image with its edited generative counterpart and incorporating fractal patterns, our approach mitigates unrealistic images and label ambiguity, improving the performance and adversarial robustness of the resulting models. Efficacy of our method is established with extensive experiments on eight public datasets for general and fine-grained classification, in both in-domain and cross-domain settings. Additionally, we demonstrate performance improvements for self-supervised learning, learning with data scarcity, and adversarial robustness. As compared to the existing state-of-the-art methods, our technique achieves stronger performance across the board.

数据增强生成模型跨域学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。