提升可见光与红外行人重识别性能,解决模态差异问题
Multi-scale Decomposed Convolution Refinement Network for Visible-Infrared Person Re-Identification

- 分层分解卷积注意力模块捕捉多尺度空间特征
- 新损失函数使跨模态特征更紧凑且可区分
- 在两个公开数据集上达到当前最佳效果
可见光-红外行人重识别(VI-ReID)面临跨模态差异大、判别能力弱的问题,现有方法在语义挖掘、跨模态融合和特征约束方面存在局限。为此,本文提出MDCRNet——一种多尺度分解卷积精炼网络,强化跨模态特征学习与判别度度量学习。具体而言,设计包含四个分层分解卷积注意力(HDCA)模块的层级学习模块(HLM),每个模块结合轻量级通道注意力与多尺度空间感知块,有效捕获多尺度空间依赖关系;同时提出联合判别度度量损失(JDML),融合新型粒度判别损失(GDL),在跨模态下同步优化类内紧凑性与类间可分性。在SYSU-MM01和RegDB数据集上的大量实验表明,MDCRNet在两个基准上均取得当前最优性能。代码已开源:https://github.com/Kevin-zms/MDCRNet。
原文摘要 · Abstract (English)
Visible-infrared person re-identification (VI-ReID) suffers from cross-modal discrepancies and limited discriminative capabilities, leading to suboptimal recognition performance. Current approaches exhibit limitations in semantic mining, cross-modal fusion and feature constraints. To tackle these challenges, we propose MDCRNet, a Multi-scale Decomposed Convolution Refinement Network that enhances cross-modal feature learning and discriminative metric learning. Specifically, we introduce a Hierarchical Learning Module (HLM) containing four Hierarchical Decomposed Convolution Attention (HDCA) modules, each equipped with lightweight channel attention and multi-scale spatial perception blocks to capture multi-scale spatial dependencies. Moreover, we develop a Joint Discriminative Metric Loss (JDML) incorporating a novel Granularity Discriminative Loss (GDL) that simultaneously optimizes intra-identity compactness and inter-identity separability across modalities. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate that MDCRNet achieves state-of-the-art performance on both benchmarks. Code is available at https://github.com/Kevin-zms/MDCRNet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。