arXiv:2606.12258cs.CV2026-06

无监督跨昼夜行人重识别,用提示词和原型学习对齐身份

Bridging Day and Night: Unsupervised Cross-Domain Re-Identification with Synergistic Prompt and Prototype Learning

论文配图:Bridging Day and Night: Unsupervised Cross-Domain Re-Identification with Synergistic Prompt and Prototype Learning
图 1 · 摘自论文原文
  • 用视觉语言模型生成无标注的文本提示,实现跨域特征对齐
  • 在无监督下达到与有监督顶尖方法相当的排名1准确率
  • 适合缺乏标注数据的夜间监控场景应用

跨域昼夜行人重识别因白天与夜间视觉差异巨大而面临挑战。现有全监督方法依赖大量人工标注,成本高且泛化能力弱。本文提出一种无监督框架,协同运用提示学习与原型表示学习,在无需人工标签的情况下关联跨域身份。第一阶段利用视觉语言模型无标注生成实例级文本提示,通过实例感知的动态偏置适配,将视觉特征与文本提示嵌入统一语义空间。第二阶段构建领域专属原型记忆库,引入两个互补模块:域内身份关联模块提升域内特征判别性,跨域原型匹配模块可靠识别正负原型对,建立昼夜间的鲁棒身份对应关系。在公开数据集上的实验验证了方法有效性,无监督设置下取得与先进有监督方法相当的Rank-1精度。

原文摘要 · Abstract (English)

Cross-domain day-night re-identification (ReID) is fundamentally challenged by the substantial visual appearance discrepancies between daytime and nighttime scenes. Existing fully supervised methods rely heavily on labor-intensive annotations, which are costly and exhibit limited generalization across domains. In this work, we investigate unsupervised day-night ReID and propose a novel framework that synergistically combines prompt learning and prototype-based representation learning to associate identities across domains without requiring manual labels. Our approach follows a progressive two-stage training strategy. In the first stage, we exploit the vision-language model to generate instance-specific textual prompts in an annotation-free manner. We employ an instance-level alignment mechanism to embed visual features and textual prompts into a unified semantic space, aligning unlabeled day/night images with learnable prompts via instance-aware dynamic-bias adaptation. In the second stage, we construct domain-specific prototype memory banks and introduce two complementary modules: i) an intra-domain identity association module to enhance feature discriminability within each domain, and ii) a cross-domain prototype matching module to reliably identify positive and negative prototype pairs, thereby establishing robust identity correspondences across day and night. Extensive experiments on public benchmarks validate the effectiveness of our method. Under the unsupervised setting, our framework attains Rank-1 accuracy comparable to state-of-the-art fully supervised methods.

行人重识别无监督学习视觉语言模型跨域匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。