统一分析扩散数据增强方法,揭示其有效性的关键因素
Diffusion-Based Data Augmentation for Image Recognition: A Systematic Analysis and Evaluation

- 将扩散数据增强拆解为微调、生成、使用三部分,构建统一分析框架
- 在多种低数据场景下对比评估,发现不同策略的优劣与适用条件
- 开源完整代码和配置,支持可复现研究和后续方法开发
基于扩散的数据增强(DiffDA)在数据稀缺条件下提升图像分类性能方面展现出巨大潜力。然而,现有研究在任务设置、模型选择和实验流程上差异显著,难以公平比较或评估其跨场景有效性。此外,对整个DiffDA流程的系统性理解仍不足。本文提出UniDiffDA,一个统一的分析框架,将DiffDA方法分解为三个核心组件:模型微调、样本生成和样本利用。这一视角使我们能够识别现有方法的关键差异,并厘清整体设计空间。基于此框架,我们建立了一个全面且公平的评估协议,在多种低数据分类任务中基准测试代表性DiffDA方法。大量实验揭示了不同策略的相对优劣与局限性,并为方法设计与部署提供了实用洞见。所有方法均在统一代码库中重新实现,代码与配置全部公开,确保可复现性并推动未来研究。
原文摘要 · Abstract (English)
Diffusion-based data augmentation (DiffDA) has emerged as a promising approach to improving classification performance under data scarcity. However, existing works vary significantly in task configurations, model choices, and experimental pipelines, making it difficult to fairly compare methods or assess their effectiveness across different scenarios. Moreover, there remains a lack of systematic understanding of the full DiffDA workflow. In this work, we introduce UniDiffDA, a unified analytical framework that decomposes DiffDA methods into three core components: model fine-tuning, sample generation, and sample utilization. This perspective enables us to identify key differences among existing methods and clarify the overall design space. Building on this framework, we develop a comprehensive and fair evaluation protocol, benchmarking representative DiffDA methods across diverse low-data classification tasks. Extensive experiments reveal the relative strengths and limitations of different DiffDA strategies and offer practical insights into method design and deployment. All methods are re-implemented within a unified codebase, with full release of code and configurations to ensure reproducibility and to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。