arXiv:2608.06943cs.CV2026-08

提出双空间一致性学习框架,提升跨模态重识别泛化能力。

Dual-Space Modality Consistency Learning for Universal Cross-Modal Re-Identification

论文配图:Dual-Space Modality Consistency Learning for Universal Cross-Modal Re-Identification
图 1 · 摘自论文原文
  • 同时在空间与频域建模模态一致性
  • 在17个评估协议上超越多个基线模型
  • 可插拔设计适用于多种跨模态场景

跨模态重识别(ReID)旨在异构成像模态间检索同一身份,广泛应用于可见光-红外行人ReID和跨模态船舶ReID。现有方法通过空间嵌入空间学习模态一致性取得良好效果,但常忽略频域差异,尤其在高频率表征上——这些特征既具判别性又对模态敏感。此外,多数方法针对特定模态设置,泛化能力受限。为此,我们提出通用跨模态ReID的双空间模态一致性学习(DSMCL)框架。DSMCL联合建模空间特征分布一致性和频域判别一致性:空间模态一致性学习(SMCL)分支采用高斯特征对齐;频域感知判别一致性学习(FDCL)策略通过身份感知的跨模态对比学习正则化高频表示。通过同时捕捉模态特异性特征与共享身份线索,DSMCL学习鲁棒表征,并构建统一框架以适配多样异构模态设置。此外,DSMCL为即插即用结构,可无缝集成至现有跨模态ReID架构。在SYSU-MM01、RegDB、LLCM、HOSS-ReID和CMShipReID上,覆盖十七个评估协议的大量实验表明,DSMCL持续提升多个代表性基线模型性能。

原文摘要 · Abstract (English)

Cross-modal Re-Identification (ReID) aims to retrieve the same identity across heterogeneous imaging modalities and has been widely studied in visible-infrared person ReID and cross-modal ship ReID. Existing methods have achieved promising performance by learning modality consistency in the spatial embedding space, yet often overlook frequency-domain modality discrepancy, particularly in high-frequency representations that are both highly discriminative and modality-sensitive. In addition, most approaches are tailored to specific modality settings, limiting their applicability across diverse cross-modal scenarios. To address these challenges, we propose a Dual-Space Modality Consistency Learning (DSMCL) framework for universal cross-modal ReID. Specifically, DSMCL jointly models spatial feature distribution consistency and frequency-domain discriminative consistency. A Spatial Modality Consistency Learning (SMCL) branch performs Gaussian-based feature alignment, while a Frequency-aware Discriminative Consistency Learning (FDCL) strategy regularizes high-frequency representations through identity-aware cross-modal contrastive learning. By jointly capturing modality-specific characteristics and modality-shared identity cues, DSMCL learns robust representations and establishes a unified framework capable of accommodating diverse heterogeneous modality settings. Moreover, DSMCL is a plug-and-play framework that can be readily integrated into existing cross-modal ReID architectures. Extensive experiments on SYSU-MM01, RegDB, LLCM, HOSS-ReID, and CMShipReID across seventeen evaluation protocols show that DSMCL consistently improves multiple representative baselines.

跨模态重识别双空间学习频域一致性通用框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。