arXiv:2510.01841cs.CV2025-10ICCV被引 2

用扩散模型先验知识提升行人搜索的定位与识别效果

Leveraging Prior Knowledge of Diffusion Model for Person Search

  • 引入扩散模型先验,分离检测与重识别特征提取
  • 在CUHK-SYSU和PRW数据集上刷新性能纪录
  • 适合关注行人搜索、跨模态特征融合的研究者

行人搜索旨在通过定位和识别查询行人,从非裁剪场景图像中完成任务。现有方法多依赖ImageNet预训练主干网络,难以捕捉复杂的空间上下文与细粒度身份线索;且共享主干特征用于检测与重识别,导致优化目标冲突。本文提出DiffPS框架,利用预训练扩散模型先验,消除两任务间的优化矛盾。分析扩散先验特性,设计三个专用模块:(i) 扩散引导区域提议网络(DGRPN)提升行人定位精度,(ii) 多尺度频域优化网络(MSFRN)缓解形状偏差,(iii) 语义自适应特征聚合网络(SFAN)利用对齐文本的扩散特征。DiffPS在CUHK-SYSU和PRW数据集上达到新最优性能。

原文摘要 · Abstract (English)

Person search aims to jointly perform person detection and re-identification by localizing and identifying a query person within a gallery of uncropped scene images. Existing methods predominantly utilize ImageNet pre-trained backbones, which may be suboptimal for capturing the complex spatial context and fine-grained identity cues necessary for person search. Moreover, they rely on a shared backbone feature for both person detection and re-identification, leading to suboptimal features due to conflicting optimization objectives. In this paper, we propose DiffPS (Diffusion Prior Knowledge for Person Search), a novel framework that leverages a pre-trained diffusion model while eliminating the optimization conflict between two sub-tasks. We analyze key properties of diffusion priors and propose three specialized modules: (i) Diffusion-Guided Region Proposal Network (DGRPN) for enhanced person localization, (ii) Multi-Scale Frequency Refinement Network (MSFRN) to mitigate shape bias, and (iii) Semantic-Adaptive Feature Aggregation Network (SFAN) to leverage text-aligned diffusion features. DiffPS sets a new state-of-the-art on CUHK-SYSU and PRW.

行人搜索扩散模型特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。