arXiv:2504.19455cs.CV2025-04被引 1

用掩码提示生成多样又风格一致的时装图像,提升少样本识别效果。

Masked Language Prompting for Generative Data Augmentation in Few-shot Fashion Style Recognition

  • 掩码参考描述中的关键词,用大语言模型补全以生成多样图像。
  • 在FashionStyle14数据集上,少样本下准确率超越基线方法。
  • 无需微调即可保持风格一致性,适合小样本时装识别场景。

时尚风格识别的数据集构建因风格概念的主观性和模糊性而困难。近年来,文本到图像模型促进了通过合成图像进行生成式数据增强,但仅依赖类别名或参考描述的方法难以平衡视觉多样性与风格一致性。本文提出一种新的提示策略——掩码语言提示(MLP),通过掩码参考描述中的选定词汇,并利用大语言模型生成语义连贯的多样化补全。该方法在保留原始描述结构语义的同时,引入与目标风格对齐的属性级变化,实现无需微调的风格一致且多样化的图像生成。在FashionStyle14数据集上的实验表明,基于MLP的增强方法在少样本条件下持续优于基于类别名和描述的基线方法,验证了其在有限监督下的有效性。

原文摘要 · Abstract (English)

Constructing dataset for fashion style recognition is challenging due to the inherent subjectivity and ambiguity of style concepts. Recent advances in text-to-image models have facilitated generative data augmentation by synthesizing images from labeled data, yet existing methods based solely on class names or reference captions often fail to balance visual diversity and style consistency. In this work, we propose \textbf{Masked Language Prompting (MLP)}, a novel prompting strategy that masks selected words in a reference caption and leverages large language models to generate diverse yet semantically coherent completions. This approach preserves the structural semantics of the original caption while introducing attribute-level variations aligned with the intended style, enabling style-consistent and diverse image generation without fine-tuning. Experimental results on the FashionStyle14 dataset demonstrate that our MLP-based augmentation consistently outperforms class-name and caption-based baselines, validating its effectiveness for fashion style recognition under limited supervision.

数据增强少样本学习风格识别提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。