arXiv:2507.15888cs.CV2025-07

生成式数据增强在物体重识别中反而降低性能,因细节丢失与域偏移。

PAT++: a cautionary tale about generative visual augmentation for Object Re-identification

  • 用扩散自蒸馏增强部分感知变压器,生成保持身份的图像。
  • 在城市元素重识别数据集上训练和查询扩展均导致性能下降。
  • 揭示生成模型在细粒度识别中的局限性,适合关注数据增强风险的研究者。

生成式数据增强在多个视觉任务中表现优异,但在物体重识别这类需保留细微视觉特征的任务中影响尚不明确。本文提出新方法PAT++,将扩散自蒸馏引入经典部分感知变压器。基于Urban Elements ReID Challenge数据集,通过生成图像进行模型训练与查询扩展的大量实验表明,性能持续下降,主因是域偏移及身份关键特征未能保留。结果挑战了生成模型可迁移至细粒度识别任务的假设,揭示当前身份保持型视觉增强方法的核心缺陷。

原文摘要 · Abstract (English)

Generative data augmentation has demonstrated gains in several vision tasks, but its impact on object re-identification - where preserving fine-grained visual details is essential - remains largely unexplored. In this work, we assess the effectiveness of identity-preserving image generation for object re-identification. Our novel pipeline, named PAT++, incorporates Diffusion Self-Distillation into the well-established Part-Aware Transformer. Using the Urban Elements ReID Challenge dataset, we conduct extensive experiments with generated images used for both model training and query expansion. Our results show consistent performance degradation, driven by domain shifts and failure to retain identity-defining features. These findings challenge assumptions about the transferability of generative models to fine-grained recognition tasks and expose key limitations in current approaches to visual augmentation for identity-preserving applications.

物体重识别生成增强扩散模型细粒度识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。