提出动态跨模态跨摄像头聚类方法,解决无监督红外可见光行人重识别中的身份分裂问题。
Dynamic Modality-Camera Invariant Clustering for Unsupervised Visible-Infrared Person Re-identification
- 融合跨模态跨摄像头距离编码,统一建模多源差异
- 动态邻域聚类使模型从判别性优化转向泛化能力提升
- 实时更新记忆库,实现模态不变表示的在线探索
无监督可见光-红外行人重识别(USL-VI-ReID)相比有监督方法更具灵活性和成本优势,受到广泛关注。现有方法仅对模态特定样本进行聚类,并采用强关联技术实现实例到聚类或聚类到聚类的跨模态关联,但忽视了跨摄像头差异,导致身份过度分裂,影响跨模态关联的准确性和可靠性。为此,本文提出一种新型动态模态-摄像头不变聚类(DMIC)框架。DMIC将模态-摄像头不变扩展(MIE)、动态邻域聚类(DNC)和混合模态对比学习(HMCL)整合为统一框架,有效消除聚类中的跨模态与跨摄像头偏差。MIE通过融合模态间与摄像头间的距离编码,在聚类层面弥合模态与摄像头差异。DNC采用两种动态搜索策略,引导网络优化目标从提升判别性转向增强跨模态与跨摄像头泛化性。HMCL用于优化实例级与聚类级分布,通过随机选取样本更新模态内与模态间训练记忆,支持模态不变表示的实时探索。大量实验表明,所提方法克服了现有聚类方法的局限,性能显著提升,大幅缩小与有监督方法的差距。
原文摘要 · Abstract (English)
Unsupervised learning visible-infrared person re-identification (USL-VI-ReID) offers a more flexible and cost-effective alternative compared to supervised methods. This field has gained increasing attention due to its promising potential. Existing methods simply cluster modality-specific samples and employ strong association techniques to achieve instance-to-cluster or cluster-to-cluster cross-modality associations. However, they ignore cross-camera differences, leading to noticeable issues with excessive splitting of identities. Consequently, this undermines the accuracy and reliability of cross-modal associations. To address these issues, we propose a novel Dynamic Modality-Camera Invariant Clustering (DMIC) framework for USL-VI-ReID. Specifically, our DMIC naturally integrates Modality-Camera Invariant Expansion (MIE), Dynamic Neighborhood Clustering (DNC) and Hybrid Modality Contrastive Learning (HMCL) into a unified framework, which eliminates both the cross-modality and cross-camera discrepancies in clustering. MIE fuses inter-modal and inter-camera distance coding to bridge the gaps between modalities and cameras at the clustering level. DNC employs two dynamic search strategies to refine the network's optimization objective, transitioning from improving discriminability to enhancing cross-modal and cross-camera generalizability. Moreover, HMCL is designed to optimize instance-level and cluster-level distributions. Memories for intra-modality and inter-modality training are updated using randomly selected samples, facilitating real-time exploration of modality-invariant representations. Extensive experiments have demonstrated that our DMIC addresses the limitations present in current clustering approaches and achieve competitive performance, which significantly reduces the performance gap with supervised methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。