提出双空间一致性学习框架,提升跨模态重识别泛化能力。
Dual-Space Modality Consistency Learning for Universal Cross-Modal Re-Identification

- 同时在空间与频域建模模态一致性
- 在17个评估协议上超越多个基线模型
- 可插拔设计适用于多种跨模态场景
跨模态重识别(ReID)旨在异构成像模态间检索同一身份,广泛应用于可见光-红外行人ReID和跨模态船舶ReID。现有方法通过空间嵌入空间学习模态一致性取得良好效果,但常忽略频域差异,尤其在高频率表征上——这些特征既具判别性又对模态敏感。此外,多数方法针对特定模态设置,泛化能力受限。为此,我们提出通用跨模态ReID的双空间模态一致性学习(DSMCL)框架。DSMCL联合建模空间特征分布一致性和频域判别一致性:空间模态一致性学习(SMCL)分支采用高斯特征对齐;频域感知判别一致性学习(FDCL)策略通过身份感知的跨模态对比学习正则化高频表示。通过同时捕捉模态特异性特征与共享身份线索,DSMCL学习鲁棒表征,并构建统一框架以适配多样异构模态设置。此外,DSMCL为即插即用结构,可无缝集成至现有跨模态ReID架构。在SYSU-MM01、RegDB、LLCM、HOSS-ReID和CMShipReID上,覆盖十七个评估协议的大量实验表明,DSMCL持续提升多个代表性基线模型性能。
原文摘要 · Abstract (English)
Cross-modal Re-Identification (ReID) aims to retrieve the same identity across heterogeneous imaging modalities and has been widely studied in visible-infrared person ReID and cross-modal ship ReID. Existing methods have achieved promising performance by learning modality consistency in the spatial embedding space, yet often overlook frequency-domain modality discrepancy, particularly in high-frequency representations that are both highly discriminative and modality-sensitive. In addition, most approaches are tailored to specific modality settings, limiting their applicability across diverse cross-modal scenarios. To address these challenges, we propose a Dual-Space Modality Consistency Learning (DSMCL) framework for universal cross-modal ReID. Specifically, DSMCL jointly models spatial feature distribution consistency and frequency-domain discriminative consistency. A Spatial Modality Consistency Learning (SMCL) branch performs Gaussian-based feature alignment, while a Frequency-aware Discriminative Consistency Learning (FDCL) strategy regularizes high-frequency representations through identity-aware cross-modal contrastive learning. By jointly capturing modality-specific characteristics and modality-shared identity cues, DSMCL learns robust representations and establishes a unified framework capable of accommodating diverse heterogeneous modality settings. Moreover, DSMCL is a plug-and-play framework that can be readily integrated into existing cross-modal ReID architectures. Extensive experiments on SYSU-MM01, RegDB, LLCM, HOSS-ReID, and CMShipReID across seventeen evaluation protocols show that DSMCL consistently improves multiple representative baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。