arXiv:2410.21318cs.CVcs.AI2024-10被引 1

通过多路径反馈调整,提升文本检索中人像匹配精度。

Multi-path Exploration and Feedback Adjustment for Text-to-Image Person Retrieval

  • 设计三路机制:内部推理、跨模态精修、判别线索修正。
  • 在三个公开数据集上超越现有方法,无需额外数据或复杂结构。
  • 适合关注跨模态检索与细粒度匹配的研究者。

基于文本的人像检索旨在通过文本描述识别特定人物。现有先进方法通常依赖视觉语言预训练(VLP)模型实现有效的跨模态对齐,但其固有的全局对齐偏差和不足的自反馈调节限制了最佳检索性能。本文提出MeFa框架,通过深度探索模内与模间内在反馈,实现针对性调整,从而获得更精确的人-文关联。具体而言,首先设计模内推理路径,生成跨模态难负样本,并利用这些样本反馈优化模内推理,增强对细微差异的敏感性;随后引入跨模态精修路径,结合全局信息与模间反馈,优化局部表征,提升全局语义表达;最后,判别线索修正路径引入次要相似性的细粒度特征作为判别线索,进一步缓解因特征差异导致的检索失败。在三个公开基准上的实验结果表明,MeFa在不需额外数据或复杂结构的情况下,实现了更优的检索性能。

原文摘要 · Abstract (English)

Text-based person retrieval aims to identify the specific persons using textual descriptions as queries. Existing ad vanced methods typically depend on vision-language pre trained (VLP) models to facilitate effective cross-modal alignment. However, the inherent constraints of VLP mod-els, which include the global alignment biases and insuffi-cient self-feedback regulation, impede optimal retrieval per formance. In this paper, we propose MeFa, a Multi-Pathway Exploration, Feedback, and Adjustment framework, which deeply explores intrinsic feedback of intra and inter-modal to make targeted adjustment, thereby achieving more precise person-text associations. Specifically, we first design an intra modal reasoning pathway that generates hard negative sam ples for cross-modal data, leveraging feedback from these samples to refine intra-modal reasoning, thereby enhancing sensitivity to subtle discrepancies. Subsequently, we intro duce a cross-modal refinement pathway that utilizes both global information and intermodal feedback to refine local in formation, thus enhancing its global semantic representation. Finally, the discriminative clue correction pathway incorpo rates fine-grained features of secondary similarity as discrim inative clues to further mitigate retrieval failures caused by disparities in these features. Experimental results on three public benchmarks demonstrate that MeFa achieves superior person retrieval performance without necessitating additional data or complex structures.

跨模态检索文本生成图像反馈机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。