提出新框架提升无监督可见光红外行人重识别性能
Extended Cross-Modality United Learning for Unsupervised Visible-Infrared Person Re-identification
- 融合跨模态聚类与实例筛选,构建更可靠的跨模态关联
- 在SYSU-MM01和RegDB上超越部分有监督方法
- 适合做无监督跨模态行人重识别的研究者参考
无监督可见光-红外行人重识别旨在从未标注的跨模态数据中学习模态不变特征并缩小模态间差异。现有方法或缺乏跨模态聚类,或过度追求聚类级关联,难以可靠学习模态不变特征。为此,本文提出扩展跨模态联合学习(ECUL)框架,包含扩展模态-相机聚类(EMCC)与两阶段记忆更新策略(TSMem)模块。ECUL自然整合了模态内聚类、跨模态聚类及跨模态实例选择,在减少噪声标签引入的同时建立紧凑准确的跨模态关联。EMCC通过扩展编码向量捕捉并过滤邻域关系,促进模态不变与相机不变知识的学习。TSMem分阶段更新记忆,为对比学习提供精确且泛化的代理点。在SYSU-MM01和RegDB数据集上的大量实验表明,所提ECUL表现优异,甚至超越某些有监督方法。
原文摘要 · Abstract (English)
Unsupervised learning visible-infrared person re-identification (USL-VI-ReID) aims to learn modality-invariant features from unlabeled cross-modality datasets and reduce the inter-modality gap. However, the existing methods lack cross-modality clustering or excessively pursue cluster-level association, which makes it difficult to perform reliable modality-invariant features learning. To deal with this issue, we propose a Extended Cross-Modality United Learning (ECUL) framework, incorporating Extended Modality-Camera Clustering (EMCC) and Two-Step Memory Updating Strategy (TSMem) modules. Specifically, we design ECUL to naturally integrates intra-modality clustering, inter-modality clustering and inter-modality instance selection, establishing compact and accurate cross-modality associations while reducing the introduction of noisy labels. Moreover, EMCC captures and filters the neighborhood relationships by extending the encoding vector, which further promotes the learning of modality-invariant and camera-invariant knowledge in terms of clustering algorithm. Finally, TSMem provides accurate and generalized proxy points for contrastive learning by updating the memory in stages. Extensive experiments results on SYSU-MM01 and RegDB datasets demonstrate that the proposed ECUL shows promising performance and even outperforms certain supervised methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。