通过不确定性感知的原型解耦,提升复杂场景下文本搜人准确率
Uncertainty-Aware Prototype Semantic Decoupling for Text-Based Person Search in Full Images
- 分粒度估计检测不确定性,动态调整搜索信心
- 在粗粒度聚类与细粒度个体层面分离学习行人原型特征
- 适合需要高精度文本搜人的实际安防场景
全图文本行人搜索(TBPS)旨在使用自然语言描述从非裁剪图像中定位目标行人。然而,在多人复杂场景中,现有方法受限于检测与匹配中的不确定性,导致性能下降。为此,我们提出UPD-TBPS框架,包含三个模块:多粒度不确定性估计(MUE)、基于原型的不确定性解耦(PUD)和跨模态重识别(ReID)。MUE通过多粒度查询识别潜在目标,并为候选分配置信度以降低早期不确定性。PUD利用视觉上下文解耦与原型挖掘,提取查询描述的行人特征,分别在粗粒度聚类层与细粒度个体层学习行人原型表示,从而减少匹配不确定性。ReID对不同置信度的候选进行评估,提升检测与检索精度。在CUHK-SYSU-TBPS与PRW-TBPS数据集上的实验验证了该框架的有效性。
原文摘要 · Abstract (English)
Text-based pedestrian search (TBPS) in full images aims to locate a target pedestrian in untrimmed images using natural language descriptions. However, in complex scenes with multiple pedestrians, existing methods are limited by uncertainties in detection and matching, leading to degraded performance. To address this, we propose UPD-TBPS, a novel framework comprising three modules: Multi-granularity Uncertainty Estimation (MUE), Prototype-based Uncertainty Decoupling (PUD), and Cross-modal Re-identification (ReID). MUE conducts multi-granularity queries to identify potential targets and assigns confidence scores to reduce early-stage uncertainty. PUD leverages visual context decoupling and prototype mining to extract features of the target pedestrian described in the query. It separates and learns pedestrian prototype representations at both the coarse-grained cluster level and the fine-grained individual level, thereby reducing matching uncertainty. ReID evaluates candidates with varying confidence levels, improving detection and retrieval accuracy. Experiments on CUHK-SYSU-TBPS and PRW-TBPS datasets validate the effectiveness of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。