arXiv:2505.06566cs.CV2025-05被引 17

解决文本搜人中的噪声数据问题,提升检索准确率。

Dynamic Uncertainty Learning with Noisy Correspondence for Text-Based Person Search

  • 用动态不确定性建模噪声,自动识别并抑制错误匹配
  • 在三个数据集上实现低噪和高噪场景下性能提升
  • 适合处理互联网收集的含噪文本图像数据

文本到图像的人体搜索旨在根据文本描述定位特定个体。为降低数据采集成本,大规模文本图像数据集通常基于网络上共现对构建,但可能引入噪声,尤其是不匹配对,导致检索性能下降。现有方法多关注负样本,反而加剧噪声影响。为此,我们提出动态不确定性与关系对齐(DURA)框架,包含关键特征选择器(KFS)和一种新损失函数——动态软门限损失(DSH-Loss)。KFS捕捉并建模噪声不确定性,提升检索可靠性;跨模态相似性的双向证据以狄利克雷分布建模,增强对噪声数据的适应性;DSH动态调节负样本难度,提升在噪声环境下的鲁棒性。在三个数据集上的实验表明,该方法具备强噪声抵抗能力,在低噪与高噪场景下均显著提升检索性能。

原文摘要 · Abstract (English)

Text-to-image person search aims to identify an individual based on a text description. To reduce data collection costs, large-scale text-image datasets are created from co-occurrence pairs found online. However, this can introduce noise, particularly mismatched pairs, which degrade retrieval performance. Existing methods often focus on negative samples, which amplify this noise. To address these issues, we propose the Dynamic Uncertainty and Relational Alignment (DURA) framework, which includes the Key Feature Selector (KFS) and a new loss function, Dynamic Softmax Hinge Loss (DSH-Loss). KFS captures and models noise uncertainty, improving retrieval reliability. The bidirectional evidence from cross-modal similarity is modeled as a Dirichlet distribution, enhancing adaptability to noisy data. DSH adjusts the difficulty of negative samples to improve robustness in noisy environments. Our experiments on three datasets show that the method offers strong noise resistance and improves retrieval performance in both low- and high-noise scenarios.

文本搜索图像检索噪声鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。