用中间图像桥接可见光与红外行人重识别,提升跨模态对齐效果。
Modality-Transition Representation Learning for Visible-Infrared Person Re-Identification
- 以生成的中间图像为桥梁,实现可见光到红外的模态过渡对齐。
- 在三个主流数据集上显著超越现有最先进方法,性能提升明显。
- 无需额外参数,推理速度不变,适合实际部署场景。
可见光-红外行人重识别(VI-ReID)技术可在背景光照变化的实际场景中关联可见光与红外模态下的行人图像。然而,两种模态间存在固有差异。现有方法多依赖中间表示来对齐同一个人的跨模态特征,这些中间表示通常通过生成中间图像(数据增强形式)或融合中间特征(增加参数量、可解释性差)获得,未能充分挖掘其价值。为此,本文提出一种基于模态转换表示学习(MTRL)的新框架,以生成的中间图像作为从可见光到红外模态的传输媒介,该中间图像与原始可见光图像完全对齐且接近红外模态特征。训练时引入模态转换对比损失和模态查询正则化损失,进一步强化跨模态特征对齐。值得注意的是,该框架不引入额外参数,推理速度与主干网络一致,同时显著提升VI-ReID性能。大量实验表明,模型在三个典型VI-ReID数据集上持续显著优于现有最先进方法。
原文摘要 · Abstract (English)
Visible-infrared person re-identification (VI-ReID) technique could associate the pedestrian images across visible and infrared modalities in the practical scenarios of background illumination changes. However, a substantial gap inherently exists between these two modalities. Besides, existing methods primarily rely on intermediate representations to align cross-modal features of the same person. The intermediate feature representations are usually create by generating intermediate images (kind of data enhancement), or fusing intermediate features (more parameters, lack of interpretability), and they do not make good use of the intermediate features. Thus, we propose a novel VI-ReID framework via Modality-Transition Representation Learning (MTRL) with a middle generated image as a transmitter from visible to infrared modals, which are fully aligned with the original visible images and similar to the infrared modality. After that, using a modality-transition contrastive loss and a modality-query regularization loss for training, which could align the cross-modal features more effectively. Notably, our proposed framework does not need any additional parameters, which achieves the same inference speed to the backbone while improving its performance on VI-ReID task. Extensive experimental results illustrate that our model significantly and consistently outperforms existing SOTAs on three typical VI-ReID datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。