arXiv:2601.11243cs.CV2026-01AAAI被引 1

无需标注数据,统一建模多场景下的人体再识别。

Image-Text Knowledge Modeling for Unsupervised Multi-Scenario Person Re-Identification

  • 用视觉语言模型构建跨场景的图像-文本知识表示。
  • 在多个场景中实现超越单场景方法的识别准确率。
  • 适合需要跨域泛化能力的研究者和工业应用。

我们提出无监督多场景人体再识别(UMS-ReID)这一新任务,旨在单一框架内拓展再识别能力以应对跨分辨率、服装变化等多种场景。为解决该问题,我们引入图像-文本知识建模(ITKM)——一个三阶段框架,充分挖掘视觉语言模型的表征能力。首先使用预训练的CLIP模型,其包含图像编码器与文本编码器。第一阶段,在图像编码器中引入场景嵌入,并微调编码器以自适应地利用多场景知识。第二阶段,优化一组学习得到的文本嵌入,使其与第一阶段生成的伪标签对齐,并引入多场景分离损失,增强不同场景间文本表示的差异性。第三阶段,设计层级异构匹配模块(包括簇级与实例级),在各场景内获取可靠异构正样本对(如同一人可见光与红外图像)。随后提出动态文本表示更新策略,维持文本与图像监督信号的一致性。在多个场景下的实验结果表明,ITKM具有显著优越性与强泛化能力,不仅优于现有场景特定方法,还通过整合多场景知识提升了整体性能。

原文摘要 · Abstract (English)

We propose unsupervised multi-scenario (UMS) person re-identification (ReID) as a new task that expands ReID across diverse scenarios (cross-resolution, clothing change, etc.) within a single coherent framework. To tackle UMS-ReID, we introduce image-text knowledge modeling (ITKM) -- a three-stage framework that effectively exploits the representational power of vision-language models. We start with a pre-trained CLIP model with an image encoder and a text encoder. In Stage I, we introduce a scenario embedding in the image encoder and fine-tune the encoder to adaptively leverage knowledge from multiple scenarios. In Stage II, we optimize a set of learned text embeddings to associate with pseudo-labels from Stage I and introduce a multi-scenario separation loss to increase the divergence between inter-scenario text representations. In Stage III, we first introduce cluster-level and instance-level heterogeneous matching modules to obtain reliable heterogeneous positive pairs (e.g., a visible image and an infrared image of the same person) within each scenario. Next, we propose a dynamic text representation update strategy to maintain consistency between text and image supervision signals. Experimental results across multiple scenarios demonstrate the superiority and generalizability of ITKM; it not only outperforms existing scenario-specific methods but also enhances overall performance by integrating knowledge from multiple scenarios.

人体再识别视觉语言模型无监督学习跨场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。