arXiv:2604.15631cs.CV2026-04被引 1

无监督视频跨模态行人重识别新方法,提升模糊标签下的身份判别能力。

Causal Bootstrapped Alignment for Unsupervised Video-Based Visible-Infrared Person Re-Identification

论文配图:Causal Bootstrapped Alignment for Unsupervised Video-Based Visible-Infrared Person Re-Identification
图 1 · 摘自论文原文
  • 利用时间一致性与跨模态一致性进行因果干预,消除伪相关性干扰。
  • 通过原型引导的不确定性精炼,解决可见光与红外模态粒度不均问题。
  • 适用于全天候监控场景,无需昂贵标注,适合实际部署应用。

VVI-ReID是实现全天候监控的关键技术,时间信息提供了超越静态图像的额外线索。然而,现有方法严重依赖需人工标注的全监督学习,限制了可扩展性。为此,本文研究无监督学习下的视频级可见光-红外行人重识别(USL-VVI-ReID),直接从无标签视频轨迹中学习身份判别表征。直接将图像级无监督方法迁移至该任务并使用通用预训练编码器会导致性能不佳:这些编码器存在身份判别弱、模态偏差强的问题,引发严重的同模态身份混淆以及可见光与红外模态间聚类粒度失衡,共同降低伪标签可靠性并阻碍跨模态对齐。为此,提出因果自举对齐(CBA)框架,显式利用视频固有先验。首先引入因果干预热身(CIW),通过时间身份一致性和跨模态身份一致性实施序列级因果干预,抑制由模态和运动引起的虚假相关性,保留身份相关语义,获得更干净的表示用于无监督聚类。其次提出原型引导的不确定性精炼(PGUR),采用粗到细对齐策略,解决跨模态粒度差异,借助可靠的可见光原型,以不确定性感知监督重新组织红外模态的欠聚类表示。在HITSZ-VCM和BUPTCampus数据集上的大量实验表明,当扩展至USL-VVI-ReID设置时,CBA显著优于现有方法。

原文摘要 · Abstract (English)

VVI-ReID is a critical technique for all-day surveillance, where temporal information provides additional cues beyond static images. However, existing approaches rely heavily on fully supervised learning with expensive cross-modality annotations, limiting scalability. To address this issue, we investigate Unsupervised Learning for VVI-ReID (USL-VVI-ReID), which learns identity-discriminative representations directly from unlabeled video tracklets. Directly extending image-based USL-VI-ReID methods to this setting with generic pretrained encoders leads to suboptimal performance. Such encoders suffer from weak identity discrimination and strong modality bias, resulting in severe intra-modality identity confusion and pronounced clustering granularity imbalance between visible and infrared modalities. These issues jointly degrade pseudo-label reliability and hinder effective cross-modality alignment. To address these challenges, we propose a Causal Bootstrapped Alignment (CBA) framework that explicitly exploits inherent video priors. First, we introduce Causal Intervention Warm-up (CIW), which performs sequence-level causal interventions by leveraging temporal identity consistency and cross-modality identity consistency to suppress modality- and motion-induced spurious correlations while preserving identity-relevant semantics, yielding cleaner representations for unsupervised clustering. Second, we propose Prototype-Guided Uncertainty Refinement (PGUR), which employs a coarse-to-fine alignment strategy to resolve cross-modality granularity mismatch, reorganizing under-clustered infrared representations under the guidance of reliable visible prototypes with uncertainty-aware supervision. Extensive experiments on the HITSZ-VCM and BUPTCampus benchmarks demonstrate that CBA significantly outperforms existing USL-VI-ReID methods when extended to the USL-VVI-ReID setting.

无监督学习视频重识别跨模态对齐因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。