arXiv:2412.08406cs.CV2024-12被引 1

让跨模态行人识别理解语义,提升匹配准确率

Embedding and Enriching Explicit Semantics for Visible-Infrared Person Re-Identification

  • 用多模型自动补全行人语言描述并对齐图文
  • 通过多视角互补增强视觉表征的语义完整性
  • 消除不同模态间噪声语义,增强跨模态一致性

可见光-红外行人重识别(VIReID)旨在跨模态检索同一身份的行人图像。现有方法仅从图像中学习视觉内容,缺乏感知高层语义的能力。本文提出嵌入与丰富显式语义(EEES)框架,学习语义丰富的跨模态行人表征。首先,借助多个大语言-视觉模型,我们设计显式语义嵌入(ESE),自动为行人补充语言描述,并将图像-文本对对齐至统一空间,从而学习与显式语义相关的视觉内容。其次,考虑到多视角信息的互补性,提出跨视角语义补偿(CVSC),构建多视角图像-文本对表示,建立其多对多匹配关系,并将知识传播至单视角表示,以补偿缺失的跨视角语义。第三,为消除不同模态间如颜色属性等噪声语义,设计跨模态语义净化(CMSP),约束跨模态图像-文本对表示间的距离接近同模态对的距离,进一步增强视觉内容的模态不变性。实验结果表明,所提方法在多个基准数据集上均显著优于现有方法。

原文摘要 · Abstract (English)

Visible-infrared person re-identification (VIReID) retrieves pedestrian images with the same identity across different modalities. Existing methods learn visual content solely from images, lacking the capability to sense high-level semantics. In this paper, we propose an Embedding and Enriching Explicit Semantics (EEES) framework to learn semantically rich cross-modality pedestrian representations. Our method offers several contributions. First, with the collaboration of multiple large language-vision models, we develop Explicit Semantics Embedding (ESE), which automatically supplements language descriptions for pedestrians and aligns image-text pairs into a common space, thereby learning visual content associated with explicit semantics. Second, recognizing the complementarity of multi-view information, we present Cross-View Semantics Compensation (CVSC), which constructs multi-view image-text pair representations, establishes their many-to-many matching, and propagates knowledge to single-view representations, thus compensating visual content with its missing cross-view semantics. Third, to eliminate noisy semantics such as conflicting color attributes in different modalities, we design Cross-Modality Semantics Purification (CMSP), which constrains the distance between inter-modality image-text pair representations to be close to that between intra-modality image-text pair representations, further enhancing the modality-invariance of visual content. Finally, experimental results demonstrate the effectiveness and superiority of the proposed EEES.

行人重识别跨模态语义增强多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。