arXiv:2509.02161cs.CV2025-09中稿 · AVSS 2025 conferen…被引 1

用扩散模型生成行人属性数据,提升零样本识别准确率。

Enhancing Zero-Shot Pedestrian Attribute Recognition with Synthetic Data Generation: A Comparative Study with Image-To-Image Diffusion Models

  • 基于图像到图像扩散模型生成行人属性合成数据。
  • 优化提示词与图像属性可使识别性能提升4.5%。
  • 适合需要增强泛化能力的智能监控研究者。

行人属性识别(PAR)旨在从图像中识别多种人体属性,广泛应用于智能监控系统。现有大规模标注数据集稀缺,导致PAR模型在遮挡、姿态变化和多样环境等复杂场景下泛化能力不足。近期扩散模型在生成多样化且逼真的合成图像方面展现出潜力,可用于扩充训练数据规模与多样性。然而,基于扩散模型的数据扩展在生成符合PAR任务需求图像方面的潜力尚未充分探索。本文系统研究了扩散模型在生成特定于PAR任务的行人合成图像中的有效性。我们识别出图像到图像扩散数据扩展的关键参数,包括文本提示、图像属性以及最新的扩散增强技术,并分析其对生成图像质量的影响。进一步地,采用表现最佳的扩展方法生成合成图像以丰富零样本训练数据,用于训练PAR模型。实验结果表明,提示词对齐与图像属性选择是图像生成的关键因素,最优配置可带来4.5%的识别性能提升。

原文摘要 · Abstract (English)

Pedestrian Attribute Recognition (PAR) involves identifying various human attributes from images with applications in intelligent monitoring systems. The scarcity of large-scale annotated datasets hinders the generalization of PAR models, specially in complex scenarios involving occlusions, varying poses, and diverse environments. Recent advances in diffusion models have shown promise for generating diverse and realistic synthetic images, allowing to expand the size and variability of training data. However, the potential of diffusion-based data expansion for generating PAR-like images remains underexplored. Such expansion may enhance the robustness and adaptability of PAR models in real-world scenarios. This paper investigates the effectiveness of diffusion models in generating synthetic pedestrian images tailored to PAR tasks. We identify key parameters of img2img diffusion-based data expansion; including text prompts, image properties, and the latest enhancements in diffusion-based data augmentation, and examine their impact on the quality of generated images for PAR. Furthermore, we employ the best-performing expansion approach to generate synthetic images for training PAR models, by enriching the zero-shot datasets. Experimental results show that prompt alignment and image properties are critical factors in image generation, with optimal selection leading to a 4.5% improvement in PAR recognition performance.

行人识别扩散模型合成数据零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。