arXiv:2412.01345cs.CV2024-12被引 9

用视觉文本对齐提升衣物变化下的身份识别准确率

See What You Seek: Semantic Contextual Integration for Cloth-Changing Person Re-Identification

  • 引入双可学习文本标记分离服装语义与干扰因素
  • 在三个数据集上超越现有方法,最高提升6.3%
  • 适合关注跨镜头身份识别的计算机视觉研究者

衣物变化的人体再识别(CC-ReID)旨在摄像头间匹配个体,尽管存在服装差异。现有方法通常缓解服装变化影响或增强身份相关特征,但难以捕捉复杂语义信息。本文提出一种新型提示学习框架——语义上下文融合(SCI),利用CLIP的视觉-文本表征能力减少服装引起的差异并强化身份线索。具体地,提出语义分离增强(SSE)模块,采用双可学习文本标记将服装相关语义与干扰因素分离,从而提取身份相关特征;同时设计语义引导交互模块(SIM),使用正交化文本特征引导视觉表示,聚焦于显著的身份特征。该语义融合提升了模型判别力,并以高维洞察丰富视觉上下文。在三个CC-ReID数据集上的大量实验表明,本方法优于当前最优技术。代码将在 https://github.com/hxy-499/CCREID-SCI 公开。

原文摘要 · Abstract (English)

Cloth-changing person re-identification (CC-ReID) aims to match individuals across surveillance cameras despite variations in clothing. Existing methods typically mitigate the impact of clothing changes or enhance identity (ID)-relevant features, but they often struggle to capture complex semantic information. In this paper, we propose a novel prompt learning framework Semantic Contextual Integration (SCI), which leverages the visual-textual representation capabilities of CLIP to reduce clothing-induced discrepancies and strengthen ID cues. Specifically, we introduce the Semantic Separation Enhancement (SSE) module, which employs dual learnable text tokens to disentangle clothing-related semantics from confounding factors, thereby isolating ID-relevant features. Furthermore, we develop a Semantic-Guided Interaction Module (SIM) that uses orthogonalized text features to guide visual representations, sharpening the focus of the model on distinctive ID characteristics. This semantic integration improves the discriminative power of the model and enriches the visual context with high-dimensional insights. Extensive experiments on three CC-ReID datasets demonstrate that our method outperforms state-of-the-art techniques. The code will be released at https://github.com/hxy-499/CCREID-SCI.

身份识别视觉语言模型提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。