通过双语义引导与全局局部互调,提升跨模态物体重识别准确率。
Multi-Modal Object Re-Identification with Dual Semantic Guidance and Global-Local Mutual Modulation

- 引入文本语义注入和掩码引导的局部-全局调制机制
- 在三个基准上显著超越现有方法,提升明显
- 适合需要高精度跨模态匹配的场景
多模态物体重识别(ReID)旨在利用不同模态间的互补信息检索目标实例。然而,现有方法面临两大挑战:一是难以利用对齐可靠的语义先验,易受背景干扰和跨模态错位影响;二是通常依赖整体特征建模,忽视全局与局部表示间的协同作用。为此,本文提出一种鲁棒的多模态 ReID 框架,包含三个核心组件:文本语义注入器(TSI)、掩码引导的全局-局部调制器(MGLM)和分层 MoE 融合模块(HMF)。TSI 通过将清晰连贯的文本特征融入视觉标记,增强语义感知。MGLM 通过软掩码与全局上下文联合引导,实现部件感知的跨模态交互,提升细粒度特征对齐。HMF 在局部语义监督下自适应聚合多光谱特征,生成判别性强且鲁棒的表示。在三个多模态 ReID 基准上的大量实验验证了该方法的有效性。代码将在接受后公开于 https://github.com/zw-absin/DSGM。
原文摘要 · Abstract (English)
Multi-modal object Re-Identification (ReID) aims to retrieve target instances by leveraging complementary information across modalities. However, existing methods suffer from two challenges. First, they often fail to exploit well-aligned and reliable semantic priors, making them vulnerable to background clutter and cross-modal misalignment. On the other hand, they typically rely on holistic feature modeling, overlooking the synergy between global and local representations. To overcome these limitations, we propose a robust multi-modal ReID framework with dual semantic guidance and global-local mutual modulation, which mainly consists of three key components, namely the Text-Semantic Injector (TSI), the Masked Global-Local Modulator (MGLM), and the Hierarchical MoE Fusion (HMF). The TSI enhances semantic awareness by integrating clean and coherent textual features into visual tokens. The MGLM enables part-aware cross-modal interaction through joint guidance from soft masks and global context, improving fine-grained feature alignment. Finally, the HMF adaptively aggregates multi-spectral features under local semantic supervision, yielding discriminative and robust representations. Extensive experiments on three multi-modal ReID benchmarks demonstrate the effectiveness of the proposed method. The code will be made publicly available at https://github.com/zw-absin/DSGM upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。