解决可见光与红外图像混杂场景下的身份识别难题,提升系统在复杂条件下的可靠性。
Dual-Edged Homogeneous-Modality Similarity: Towards Visible-Infrared Modality-Incomplete Person Re-Identification with Modality Adaptive Matching

- 设计可自适应匹配的Transformer模型,分离模态特异性与共享特征
- 在新构建的双模态不完整数据集上实现显著性能提升
- 适合需要应对真实世界多模态不确定性的视觉识别应用
可见光-红外行人重识别(VI-ReID)通常假设查询与图库来自异质模态,但在开放世界中,两者均可能包含同质与异质模态图像。查询可能为仅可见光、仅红外或混合模态,图库则因长期采集呈现多模态。现有方法基于异质模态检索范式,面临三大可信性挑战:同模态相似性导致匹配冲突、模态不确定性干扰、未知模态组合引发鲁棒性下降。为此,我们提出可见光-红外模态不完整重识别(VIMI-ReID)任务,重构现有数据集构建SYSU-VIMI与RegDB-VIMI基准。不可预测的模态组合及同模态样本内在相似性使现有方法性能大幅下降。我们提出模态自适应匹配变换器(MAMT),包含发散变换器模块(DTM)和共享变换器模块(STM),分别提取模态特异与共享特征。通过发散损失引导,DTM增强同模态内的判别性;模态自适应匹配模块(MAM)根据查询-图库模态关系动态融合特征,在任意且不确定的模态条件下实现稳定匹配。大量实验验证了MAMT的有效性与适应性。
原文摘要 · Abstract (English)
Visible-Infrared Person Re-Identification (VI-ReID) operates under a closed-world assumption, where queries and galleries are from heterogeneous modalities. However, in open-world scenarios, both sets are likely to contain homogeneous and heterogeneous modality images. A query may consist of visible-only, infrared-only, or mixed-modality images, while galleries present multi-modal images over long-term collection. Under these conditions, VI-ReID methods, built on a heterogeneous-modality retrieval paradigm, suffer from three trustworthiness challenges: matching conflicts due to high homogeneous-modality similarity, interference from modality uncertainty, and robustness degradation induced by unknown modality combinations. They fail to meet the requirements of trustworthy visual recognition in reliability, consistency, and dynamic adaptability. To address these challenges, we formalize the Visible-Infrared Modality-Incomplete Re-Identification (VIMI-ReID) task. We reorganize existing datasets to construct the SYSU-VIMI and RegDB-VIMI benchmarks. The unpredictable modality combinations and inherent similarity of homogeneous-modality samples in VIMI-ReID cause a significant performance drop in existing VI-ReID methods. We propose the Modality Adaptive Matching Transformer (MAMT). It employs a Divergence Transformer Module (DTM) and a Shared Transformer Module (STM) to extract modality-specific and modality-shared features, respectively. Guided by a divergence loss, the DTM enriches modality-specific features with modality-style information to enhance discriminability within the same modality. A Modality Adaptive Matching Module (MAM) dynamically fuses features according to the query-gallery modality relationship, enabling stable matching under arbitrary and uncertain modality conditions. Extensive experiments on the VIMI benchmarks demonstrate the effectiveness and adaptability of MAMT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。