首次系统验证纯合成数据在文本行人检索中的可行性。
An Empirical Study of Validating Synthetic Data for Text-Based Person Retrieval
- 提出无需真实数据的统一合成流水线,自动构建图像与描述。
- 实验证明合成数据可独立使用或补充真实数据,效果稳定。
- 适合隐私敏感场景或标注成本高的研究者参考。
数据在文本行人检索(TBPR)研究中起关键作用。主流范式依赖真实人物图像与人工文本标注进行训练,存在隐私风险和标注负担。已有研究尝试生成合成数据,但仍以真实数据为基础,延续相同局限。纯合成TBPR数据的可行性尚未探索,且缺乏对不同现实场景下合成数据有效性的系统研究。本文首次开展针对TBPR的合成数据全面实证研究,包含两方面:(1) 提出一种完全不依赖真实人像的统一数据合成流程,结合跨类别图像生成模块(通过自动提示构造策略生成多样化身份图像)与跨类别增强模块(通过文本驱动图像编辑提升身份多样性);(2) 借助该流程与自动文本描述生成,在多种场景下开展大规模实验,揭示合成数据作为独立替代或真实数据补充的实际效用边界。
原文摘要 · Abstract (English)
Data plays a pivotal role in Text-Based Person Retrieval (TBPR) research. Mainstream research paradigm necessitates real-world person images with manual textual annotations for training models, posing privacy concerns and annotation burdens. Several pioneering efforts explore synthetic data generation, and yet still depend on real data as a foundation, inheriting the same limitations. The feasibility of purely synthetic TBPR data remains unexplored, and there is currently no systematic study on the effectiveness boundaries of synthetic data across various real-world scenarios. In this work, we present the first comprehensive empirical study of synthetic data for TBPR, with two key aspects. (1) We propose a unified data synthesis pipeline that can operate entirely without real person data. It combines an inter-class image generation module that produces diverse identity-centric images by means of an automatic prompt construction strategy, and an intra-class augmentation module that enhances identity variation through text-driven image editing. (2) Leveraging this pipeline and an automatic textual description generation, we explore the effectiveness of synthetic data in diverse scenarios through extensive experiments, to reveal its practical utility as either a standalone replacement or a complementary augmentation to real data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。