arXiv:2501.17586cs.CVcs.LG2025-01

通过动态加权难样本提升文本行人检索效果

Boosting Weak Positives for Text Based Person Search

  • 用自适应权重增强难匹配的正样本
  • 在四个数据集上均实现性能提升
  • 适合关注细粒度跨模态检索的研究者

大规模视觉语言模型虽已革新跨模态检索,但文本基于行人搜索(TBPS)因数据稀缺和任务细粒度仍具挑战。现有方法多聚焦于对齐图像-文本对至统一表示空间,却忽视真实正样本间相似度差异。这导致模型偏向简单样本,部分近期方法甚至将难样本当作噪声丢弃。本文提出一种增强技术,动态识别并强化训练中的难样本。受经典提升算法启发,我们动态调整弱正样本的权重,即那些排名首位却不匹配查询身份的样本。该权重使这些误排序样本对损失函数贡献更大,迫使网络更关注此类困难样本。所提模块在四个行人数据集上均取得性能提升,验证了其有效性。

原文摘要 · Abstract (English)

Large vision-language models have revolutionized cross-modal object retrieval, but text-based person search (TBPS) remains a challenging task due to limited data and fine-grained nature of the task. Existing methods primarily focus on aligning image-text pairs into a common representation space, often disregarding the fact that real world positive image-text pairs share a varied degree of similarity in between them. This leads models to prioritize easy pairs, and in some recent approaches, challenging samples are discarded as noise during training. In this work, we introduce a boosting technique that dynamically identifies and emphasizes these challenging samples during training. Our approach is motivated from classical boosting technique and dynamically updates the weights of the weak positives, wherein, the rank-1 match does not share the identity of the query. The weight allows these misranked pairs to contribute more towards the loss and the network has to pay more attention towards such samples. Our method achieves improved performance across four pedestrian datasets, demonstrating the effectiveness of our proposed module.

文本检索行人搜索跨模态难样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。