提出AMINet模型,提升可见光与红外行人重识别的光照不变性与跨模态对齐能力。
Adaptive Illumination-Invariant Synergistic Feature Integration in a Stratified Granular Framework for Visible-Infrared Re-Identification
- 分层粒度特征提取+自适应交互融合,增强遮挡与背景干扰下的识别鲁棒性。
- 在SYSU-MM01数据集上达74.75% Rank-1准确率,优于当前最优方法3.95%。
- 适合夜间监控、搜救等复杂光照场景下的跨模态行人识别任务。
可见光-红外行人重识别(VI-ReID)在搜救、基础设施保护和夜间监控中具有重要作用,但受模态差异、光照变化和频繁遮挡影响显著。为此,我们提出自适应模态交互网络AMINet。该模型采用多粒度特征提取,从全身与上半身图像中捕获完整身份属性,提升对遮挡和背景杂乱的鲁棒性;通过交互式特征融合策略实现深度的模内与跨模态对齐,增强泛化能力并有效弥合RGB-IR模态差距。此外,利用相位一致性实现光照不变特征提取,并引入自适应多尺度核MMD对不同尺度下的特征分布进行对齐。在基准数据集上的大量实验表明,该方法在SYSU-MM01上达到74.75%的Rank-1准确率,较基线提升7.93%,超越当前最先进方法3.95%。
原文摘要 · Abstract (English)
Visible-Infrared Person Re-Identification (VI-ReID) plays a crucial role in applications such as search and rescue, infrastructure protection, and nighttime surveillance. However, it faces significant challenges due to modality discrepancies, varying illumination, and frequent occlusions. To overcome these obstacles, we propose \textbf{AMINet}, an Adaptive Modality Interaction Network. AMINet employs multi-granularity feature extraction to capture comprehensive identity attributes from both full-body and upper-body images, improving robustness against occlusions and background clutter. The model integrates an interactive feature fusion strategy for deep intra-modal and cross-modal alignment, enhancing generalization and effectively bridging the RGB-IR modality gap. Furthermore, AMINet utilizes phase congruency for robust, illumination-invariant feature extraction and incorporates an adaptive multi-scale kernel MMD to align feature distributions across varying scales. Extensive experiments on benchmark datasets demonstrate the effectiveness of our approach, achieving a Rank-1 accuracy of $74.75\%$ on SYSU-MM01, surpassing the baseline by $7.93\%$ and outperforming the current state-of-the-art by $3.95\%$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。