arXiv:2509.02099cs.CV2025-09被引 1

用文本驱动的生成模型扩充数据,提升行人属性识别效果。

A Data-Centric Approach to Pedestrian Attribute Recognition: Synthetic Augmentation via Prompt-driven Diffusion Models

  • 基于文本提示生成合成行人图像,保持数据集一致性。
  • 显著提升低频属性识别率,整体性能超越原有模型。
  • 无需修改模型结构,适合实际场景快速部署。

行人属性识别(PAR)因真实数据中属性种类繁多而极具挑战性。传统方法依赖复杂模型,但性能常受限于训练数据的不足,尤其是某些属性样本稀少。本文提出一种以数据为中心的改进方案:通过文本提示驱动的扩散模型生成合成行人图像,实现数据增强。首先,定义协议识别多个数据集中识别率低的属性;其次,设计提示驱动的生成流程,在保持PAR数据集一致性的同时生成合成图像;最后,提出将合成样本融入训练的策略,结合提示标注规则并调整损失函数。在主流PAR数据集上的实验表明,该方法不仅显著提升稀有属性的识别效果,还全面优化了整体性能。尤为关键的是,该方法在不改变模型架构的前提下增强了零样本泛化能力,为现实世界中的行人属性识别提供了高效且可扩展的解决方案。

原文摘要 · Abstract (English)

Pedestrian Attribute Recognition (PAR) is a challenging task as models are required to generalize across numerous attributes in real-world data. Traditional approaches focus on complex methods, yet recognition performance is often constrained by training dataset limitations, particularly the under-representation of certain attributes. In this paper, we propose a data-centric approach to improve PAR by synthetic data augmentation guided by textual descriptions. First, we define a protocol to identify weakly recognized attributes across multiple datasets. Second, we propose a prompt-driven pipeline that leverages diffusion models to generate synthetic pedestrian images while preserving the consistency of PAR datasets. Finally, we derive a strategy to seamlessly incorporate synthetic samples into training data, which considers prompt-based annotation rules and modifies the loss function. Results on popular PAR datasets demonstrate that our approach not only boosts recognition of underrepresented attributes but also improves overall model performance beyond the targeted attributes. Notably, this approach strengthens zero-shot generalization without requiring architectural changes of the model, presenting an efficient and scalable solution to improve the recognition of attributes of pedestrians in the real world.

行人属性数据增强扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。