无需标签,通过时间原型实现跨模态视频行人重识别
Temporal Prototyping and Hierarchical Alignment for Unsupervised Video-based Visible-Infrared Person Re-Identification

- 构建分层时间原型,利用视频时序结构学习无监督特征
- 在HITSZ-VCM和BUPTCampus上达到当前最优无监督性能
- 适合实际部署中缺乏标注数据的全天候监控场景
可见光-红外行人重识别(VI-ReID)可实现全天候监控下的跨模态身份匹配,但现有方法多聚焦于图像级或依赖昂贵的身份标注。虽然视频级VI-ReID已出现以利用时序动态提升鲁棒性,但多数研究仍局限于有监督设置。关键问题是:在无身份标签条件下,仅从RGB与红外轨迹片段学习的无监督视频级VI-ReID尚未被充分探索,却具有重要现实意义。为此,我们提出HiTPro(分层时间原型)框架,一种无需显式硬伪标签分配的原型驱动方法。首先通过时序感知特征编码器提取帧级判别特征并聚合为轨迹级表示;随后基于同相机子轨迹聚类构建可靠原型;再通过分层跨原型对齐,分两阶段进行正样本挖掘:从同模态关联逐步推进至跨模态匹配,结合动态阈值策略与软权重分配;最后采用分层对比学习,在三个层次上优化特征与原型对齐:同相机判别、跨相机同模态一致性、跨模态不变性。在HITSZ-VCM和BUPTCampus数据集上的大量实验表明,HiTPro在全无监督设置下达到最先进性能,显著优于适配基线,为后续研究建立了强基准。
原文摘要 · Abstract (English)
Visible-infrared person re-identification (VI-ReID) enables cross-modality identity matching for all-day surveillance, yet existing methods predominantly focus on the image level or rely heavily on costly identity annotations. While video-based VI-ReID has recently emerged to exploit temporal dynamics for improved robustness, existing studies remain limited to supervised settings. Crucially, the unsupervised video VI-ReID problem, where models must learn from RGB and infrared tracklets without identity labels, remains largely unexplored despite its practical importance in real-world deployment. To bridge this gap, we propose HiTPro (Hierarchical Temporal Prototyping), a prototype-driven framework without explicit hard pseudo-label assignment for unsupervised video-based VI-ReID. HiTPro begins with an efficient Temporal-aware Feature Encoder that first extracts discriminative frame-level features and then aggregates them into a robust tracklet-level representation. Building upon these features, HiTPro first constructs reliable intra-camera prototypes via Intra-Camera Tracklet Prototyping by aggregating features from temporally partitioned sub-tracklets. Through Hierarchical Cross-Prototype Alignment, we perform a two-stage positive mining process: progressing from within-modality associations to cross-modality matching, enhanced by Dynamic Threshold Strategy and Soft Weight Assignment. Finally, {Hierarchical Contrastive Learning} progressively optimizes feature-prototype alignment across three levels: intra-camera discrimination, cross-camera same-modality consistency, and cross-modality invariance. Extensive experiments on HITSZ-VCM and BUPTCampus demonstrate that HiTPro achieves state-of-the-art performance under fully unsupervised settings, significantly outperforming adapted baselines and establishes a strong baseline for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。