arXiv:2504.04510cs.CV2025-04被引 1

用带属性的提示词生成更丰富的合成图像,提升零样本领域图像分类性能

Attributed Synthetic Data Generation for Zero-shot Domain-specific Image Classification

  • 利用大语言模型生成带属性的提示词,增强合成图像多样性
  • 在两个细粒度数据集上,合成训练使分类准确率显著超越CLIP零样本和简单提示策略
  • 适合需要零样本训练但缺乏真实标注数据的研究者

零样本领域特定图像分类在无真实域内训练样本情况下极具挑战。近期研究通过文本到图像模型利用文本知识生成域内训练图像。然而,现有方法高度依赖简单提示策略,限制了合成图像的多样性,导致性能低于真实图像。本文提出AttrSyn,借助大语言模型生成带属性的提示词,以生成更具多样性的属性化合成图像。在两个细粒度数据集上的零样本领域特定图像分类实验表明,使用AttrSyn生成的合成图像进行训练,在多数情况下显著优于CLIP的零样本分类,并持续超越简单提示策略。

原文摘要 · Abstract (English)

Zero-shot domain-specific image classification is challenging in classifying real images without ground-truth in-domain training examples. Recent research involved knowledge from texts with a text-to-image model to generate in-domain training images in zero-shot scenarios. However, existing methods heavily rely on simple prompt strategies, limiting the diversity of synthetic training images, thus leading to inferior performance compared to real images. In this paper, we propose AttrSyn, which leverages large language models to generate attributed prompts. These prompts allow for the generation of more diverse attributed synthetic images. Experiments for zero-shot domain-specific image classification on two fine-grained datasets show that training with synthetic images generated by AttrSyn significantly outperforms CLIP's zero-shot classification under most situations and consistently surpasses simple prompt strategies.

零样本学习合成数据图像生成属性生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。