arXiv:2505.11813cs.CVcs.AI2025-05被引 4

用注意力引导混合与微调扩散模型,提升图像分类的数据增强效果。

SGD-Mix: Enhancing Domain-Specific Image Classification with Label-Preserving Data Augmentation

  • 通过显著性引导混合和微调扩散模型,兼顾数据多样性、真实性与标签清晰度。
  • 在细粒度、长尾、少样本等任务上均超越现有最优方法。
  • 适合需要高质量数据增强的领域特定图像分类场景。

针对领域特定图像分类中的数据增强难题,现有基于生成扩散模型的方法难以同时兼顾数据多样性、真实性和标签清晰度,且常忽视扩散模型对模型特性敏感及强变换下随机性强的问题。本文提出SGD-Mix框架,将多样性、真实性和标签一致性显式融入增强流程。通过显著性引导的混合策略与微调后的扩散模型,有效保留前景语义、丰富背景多样性并确保标签一致性,同时缓解扩散模型固有缺陷。在细粒度、长尾、少样本及背景鲁棒性任务上的大量实验表明,该方法显著优于当前最先进方法。

原文摘要 · Abstract (English)

Data augmentation for domain-specific image classification tasks often struggles to simultaneously address diversity, faithfulness, and label clarity of generated data, leading to suboptimal performance in downstream tasks. While existing generative diffusion model-based methods aim to enhance augmentation, they fail to cohesively tackle these three critical aspects and often overlook intrinsic challenges of diffusion models, such as sensitivity to model characteristics and stochasticity under strong transformations. In this paper, we propose a novel framework that explicitly integrates diversity, faithfulness, and label clarity into the augmentation process. Our approach employs saliency-guided mixing and a fine-tuned diffusion model to preserve foreground semantics, enrich background diversity, and ensure label consistency, while mitigating diffusion model limitations. Extensive experiments across fine-grained, long-tail, few-shot, and background robustness tasks demonstrate our method's superior performance over state-of-the-art approaches.

图像增强扩散模型分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。