提出频域统一框架,提升多模态行人重识别的细节刻画能力
FUSE: Frequency-domain Unification and Spectral Energy Alignment for Multi-modal Object Re-Identification

- 分频建模:将特征分解为低、中、高频子空间,实现层级化频域表示
- 跨模态对齐:通过能量一致性和子空间互补性提升不同模态间匹配精度
- 适应性强:引入可学习频域调制,增强光照和传感器差异下的鲁棒性
尽管多模态行人重识别(ReID)取得显著进展,现有方法普遍侧重低频信息,关注颜色、光照和粗略外观等属性,忽视了编码几何、纹理及身份判别细节的中高频结构。这种不平衡导致频谱表征不完整,跨模态对齐不稳定。为此,本文提出FUSE,一个基于频域的框架,将多模态ReID重构为频谱解耦与能量对齐的两阶段过程。提出的频谱分解模块(SDM)自适应地将特征划分为低、中、高频子空间,支持层级化频域建模。跨模态对齐模块(CAM)通过频率一致性正则化,强制跨模态间的能量对齐与子空间互补。此外,FUSE引入可学习频域调制,增强在不同光照和异构传感器条件下的鲁棒性。在RGBNT201、RGBNT100和MSVR310数据集上的大量实验表明,FUSE实现了9.1%的mAP和9.5%的Rank-1性能提升,建立了一个可解释的多模态表示学习频域范式。
原文摘要 · Abstract (English)
Despite significant progress in multi-modal Re-Identification (ReID), existing methods tend to emphasize low-frequency cues. Consequently, they focus on attributes such as color, illumination, and coarse appearance, while overlooking mid and high-frequency structures that encode geometric, textural, and identity-discriminative details. This imbalance leads to incomplete spectral representations and unstable cross-modal alignment. To overcome these limitations, we introduce FUSE, a frequency-domain framework that reformulates multi-modal ReID as a two-stage process of spectral disentanglement and energy alignment. The proposed Spectral Decomposition Module (SDM) adaptively partitions features into low, mid, and high-frequency subspaces, enabling hierarchical spectral modeling. The Cross-Modal Alignment Module (CAM) further enforces energy alignment and subspace complementarity across modalities via frequency-consistency regularization. In addition, FUSE incorporates learnable frequency modulation to enhance robustness under varying illumination and heterogeneous sensor conditions. Extensive experiments on RGBNT201, RGBNT100, and MSVR310 show that FUSE achieves 9.1\% mAP and 9.5\% Rank-1 improvements, establishing an interpretable frequency-domain paradigm for multi-modal representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。