arXiv:2504.19244cs.CV2025-04IJCV被引 18

通过细粒度语义对齐与协同优化,提升无监督跨模态行人重识别性能

Semantic-Aligned Learning with Collaborative Refinement for Unsupervised VI-ReID

  • 双向伪标签统一+细粒度语义对齐,增强跨模态特征一致性
  • 在SYSU-MM01和CUHK03数据集上,mAP分别达64.2%和73.8%,优于现有方法
  • 适合关注无监督跨模态学习、细粒度特征对齐的研究者

无监督可见光-红外行人重识别(USL-VI-ReID)旨在无需人工标注的情况下,匹配不同模态下同一行人的图像。现有方法通过标签关联算法统一跨模态伪标签,并设计对比学习框架进行全局特征学习,但忽略了由细粒度模式带来的跨模态特征表示与伪标签分布差异,导致模态共享学习不足。为此,本文提出语义对齐协同精炼框架(SALCR),针对各模态强调的特定细粒度模式构建优化目标,实现模态间伪标签分布互补对齐。首先引入双路全局学习伪标签统一模块(DAGI),实现跨模态实例的双向伪标签融合;随后采用细粒度语义对齐学习模块(FGSAL),从跨模态实例中挖掘各模态强调的部件级语义对齐模式,并基于对齐特征及其标签空间制定优化目标。为缓解噪声伪标签影响,设计全局-部件协同精炼模块(GPCR),动态挖掘可靠正样本集并优化实例间关系。大量实验表明,该方法在SYSU-MM01和CUHK03数据集上分别取得64.2%和73.8%的mAP,显著优于当前最先进方法。

原文摘要 · Abstract (English)

Unsupervised visible-infrared person re-identification (USL-VI-ReID) seeks to match pedestrian images of the same individual across different modalities without human annotations for model learning. Previous methods unify pseudo-labels of cross-modality images through label association algorithms and then design contrastive learning framework for global feature learning. However, these methods overlook the cross-modality variations in feature representation and pseudo-label distributions brought by fine-grained patterns. This insight results in insufficient modality-shared learning when only global features are optimized. To address this issue, we propose a Semantic-Aligned Learning with Collaborative Refinement (SALCR) framework, which builds up optimization objective for specific fine-grained patterns emphasized by each modality, thereby achieving complementary alignment between the label distributions of different modalities. Specifically, we first introduce a Dual Association with Global Learning (DAGI) module to unify the pseudo-labels of cross-modality instances in a bi-directional manner. Afterward, a Fine-Grained Semantic-Aligned Learning (FGSAL) module is carried out to explore part-level semantic-aligned patterns emphasized by each modality from cross-modality instances. Optimization objective is then formulated based on the semantic-aligned features and their corresponding label space. To alleviate the side-effects arising from noisy pseudo-labels, we propose a Global-Part Collaborative Refinement (GPCR) module to mine reliable positive sample sets for the global and part features dynamically and optimize the inter-instance relationships. Extensive experiments demonstrate the effectiveness of the proposed method, which achieves superior performances to state-of-the-art methods. Our code is available at \href{https://github.com/FranklinLingfeng/code-for-SALCR}.

行人重识别无监督学习跨模态细粒度对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。