通过部件级跨模态对齐,提升文本搜人精度。
Improving Text-based Person Search via Part-level Cross-modal Correspondence
- 构建粗到细的编码解码模型,无监督实现图文语义对齐。
- 提出基于共性度量的排序损失,增强细粒度部件学习。
- 在三个公开数据集上刷新最佳性能,适合多模态检索研究者。
文本基行人搜索旨在根据自然语言描述定位最相关的行人图像。该任务的核心挑战在于图像与文本间存在巨大语义鸿沟,导致难以建立准确对应并区分个体间的细微差异。为此,我们提出一种高效的编码器-解码器模型,能够无监督地生成粗粒度到细粒度的嵌入向量,在跨模态间实现语义对齐。另一个挑战是仅以行人ID为监督信号时,如何捕捉细粒度信息——因缺乏部件级标注,不同个体的相似身体部位被视作不同。为此,我们设计了一种新型排序损失,称为基于共性的边界排序损失(commonality-based margin ranking loss),量化每个身体部件的共性程度,并在学习过程中体现该信息。结果表明,该方法在三个公开基准上均取得最优表现。
原文摘要 · Abstract (English)
Text-based person search is the task of finding person images that are the most relevant to the natural language text description given as query. The main challenge of this task is a large gap between the target images and text queries, which makes it difficult to establish correspondence and distinguish subtle differences across people. To address this challenge, we introduce an efficient encoder-decoder model that extracts coarse-to-fine embedding vectors which are semantically aligned across the two modalities without supervision for the alignment. There is another challenge of learning to capture fine-grained information with only person IDs as supervision, where similar body parts of different individuals are considered different due to the lack of part-level supervision. To tackle this, we propose a novel ranking loss, dubbed commonality-based margin ranking loss, which quantifies the degree of commonality of each body part and reflects it during the learning of fine-grained body part details. As a consequence, it enables our method to achieve the best records on three public benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。