零样本交互式行人检索框架,提升开放世界下的语义理解与跨场景适应能力。
FitPro: A Zero-Shot Framework for Interactive Text-based Pedestrian Retrieval in Open World
- 通过提示引导对比解码生成高质量行人描述,缓解零样本下的语义漂移。
- 多视角增量语义挖掘构建全局表征,增强对视角变化和细粒度描述的鲁棒性。
- 按查询类型动态优化检索流程,适合多模态、多视角输入的实际应用。
文本驱动行人检索(TPR)旨在根据自然语言描述从视觉场景中定位特定行人。尽管现有方法在受限设置下取得进展,开放世界中的交互式检索仍面临模型泛化能力弱和语义理解不足的问题。为此,我们提出FitPro,一个具备增强语义理解与跨场景适应性的开放世界交互式零样本TPR框架。该框架包含三项创新组件:特征对比解码(FCD)、增量语义挖掘(ISM)和查询感知层级检索(QHR)。FCD利用提示引导的对比解码,从去噪图像生成高质量结构化行人描述,有效缓解零样本场景下的语义漂移。ISM通过多视角观测构建整体行人表征,实现多轮交互中的全局语义建模,提升对视角变化和细粒度描述差异的鲁棒性。QHR根据查询类型动态优化检索流程,支持对多模态与多视角输入的高效适配。在五个公开数据集和两种评估协议上的大量实验表明,FitPro显著突破了现有方法在交互式检索中的泛化局限与语义建模瓶颈,为实际部署铺平道路。
原文摘要 · Abstract (English)
Text-based Pedestrian Retrieval (TPR) deals with retrieving specific target pedestrians in visual scenes according to natural language descriptions. Although existing methods have achieved progress under constrained settings, interactive retrieval in the open-world scenario still suffers from limited model generalization and insufficient semantic understanding. To address these challenges, we propose FitPro, an open-world interactive zero-shot TPR framework with enhanced semantic comprehension and cross-scene adaptability. FitPro has three innovative components: Feature Contrastive Decoding (FCD), Incremental Semantic Mining (ISM), and Query-aware Hierarchical Retrieval (QHR). The FCD integrates prompt-guided contrastive decoding to generate high-quality structured pedestrian descriptions from denoised images, effectively alleviating semantic drift in zero-shot scenarios. The ISM constructs holistic pedestrian representations from multi-view observations to achieve global semantic modeling in multi-turn interactions, thereby improving robustness against viewpoint shifts and fine-grained variations in descriptions. The QHR dynamically optimizes the retrieval pipeline according to query types, enabling efficient adaptation to multi-modal and multi-view inputs. Extensive experiments on five public datasets and two evaluation protocols demonstrate that FitPro significantly overcomes the generalization limitations and semantic modeling constraints of existing methods in interactive retrieval, paving the way for practical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。