arXiv:2604.08877cs.CV2026-04

利用弱正样本不确定性提升文本搜索行人效果

Harnessing Weak Pair Uncertainty for Text-based Person Search

  • 通过估计图像-文本对的置信度,动态调整损失权重
  • 在三个数据集上分别提升3.06%、3.55%、6.94%的mAP
  • 特别适合处理多视角描述同一人的弱正样本场景

本文研究基于文本的行人搜索任务,旨在通过自然语言描述检索目标行人。现有方法通常依赖视觉与文本模态间严格的一一对应匹配(如对比学习),但忽略了同一人不同视角(摄像头)生成的弱正样本对。为充分挖掘弱正样本,本文提出一种不确定性感知方法,显式估计图像-文本对的不确定性,并以平滑方式将其融入优化过程。方法包含两个模块:不确定性估计模块用于获取正样本对的相对置信度;不确定性正则化模块则根据预测不确定性自适应调整损失权重。此外,引入分组图像-文本匹配损失,进一步促进弱正样本间的表示空间对齐。相比现有方法,该方法有效避免模型排除潜在的弱正候选。在CUHK-PEDES、RSTPReid和ICFG-PEDES三个常用数据集上的实验表明,mAP分别提升+3.06%、+3.55%和+6.94%。

原文摘要 · Abstract (English)

In this paper, we study the text-based person search, which is to retrieve the person of interest via natural language description. Prevailing methods usually focus on the strict one-to-one correspondence pair matching between the visual and textual modality, such as contrastive learning. However, such a paradigm unintentionally disregards the weak positive image-text pairs, which are of the same person but the text descriptions are annotated from different views (cameras). To take full use of weak positives, we introduce an uncertainty-aware method to explicitly estimate image-text pair uncertainty, and incorporate the uncertainty into the optimization procedure in a smooth manner. Specifically, our method contains two modules: uncertainty estimation and uncertainty regularization. (1) Uncertainty estimation is to obtain the relative confidence on the given positive pairs; (2) Based on the predicted uncertainty, we propose the uncertainty regularization to adaptively adjust loss weight. Additionally, we introduce a group-wise image-text matching loss to further facilitate the representation space among the weak pairs. Compared with existing methods, the proposed method explicitly prevents the model from pushing away potentially weak positive candidates. Extensive experiments on three widely-used datasets, .e.g, CUHK-PEDES, RSTPReid and ICFG-PEDES, verify the mAP improvement of our method against existing competitive methods +3.06%, +3.55% and +6.94%, respectively.

文本搜索行人重识别弱监督不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。