arXiv:2509.22635cs.CVcs.LG2025-09被引 3

无需训练即可生成判别性强的合成图像,提升少样本分类效果

Training-Free Synthetic Data Generation with Dual IP-Adapter Guidance

  • 用双提示词控制正负样本图像生成,实现精准特征引导
  • 在10个基准数据集上达到顶尖或相当水平,无需微调
  • 完全免外部工具,适合细粒度分类任务快速部署

少样本图像分类因标注样本有限而面临挑战。现有方法虽尝试利用文本到图像扩散模型生成合成训练数据,但常需大量模型微调或依赖外部信息源。我们提出一种全新的无训练方法DIPSY,通过IP-Adapter实现图像到图像转换,仅使用少量已知样本即可生成高度判别性的合成图像。DIPSY引入三项关键创新:(1) 扩展的无分类器引导机制,可独立控制正负图像条件;(2) 基于类别相似性的采样策略,识别有效的对比样本;(3) 简单高效的流水线,无需模型微调或外部描述生成与图像过滤。在十个基准数据集上的实验表明,该方法性能达到当前最优或相当水平,同时避免了生成模型适配及对外部工具的依赖。结果表明,结合正负提示的双重图像引导在生成类判别特征方面尤为有效,尤其适用于细粒度分类任务。

原文摘要 · Abstract (English)

Few-shot image classification remains challenging due to the limited availability of labeled examples. Recent approaches have explored generating synthetic training data using text-to-image diffusion models, but often require extensive model fine-tuning or external information sources. We present a novel training-free approach, called DIPSY, that leverages IP-Adapter for image-to-image translation to generate highly discriminative synthetic images using only the available few-shot examples. DIPSY introduces three key innovations: (1) an extended classifier-free guidance scheme that enables independent control over positive and negative image conditioning; (2) a class similarity-based sampling strategy that identifies effective contrastive examples; and (3) a simple yet effective pipeline that requires no model fine-tuning or external captioning and filtering. Experiments across ten benchmark datasets demonstrate that our approach achieves state-of-the-art or comparable performance, while eliminating the need for generative model adaptation or reliance on external tools for caption generation and image filtering. Our results highlight the effectiveness of leveraging dual image prompting with positive-negative guidance for generating class-discriminative features, particularly for fine-grained classification tasks.

少样本学习合成数据扩散模型图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。