通过分割引导与跨模态超图交互,提升多模态物体重识别准确率。
STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification

- 用分割掩码引导特征调制,增强前景、抑制背景噪声。
- 保留全部令牌并自适应重分配,提取紧凑有信息量的表示。
- 构建跨模态超图捕捉高级语义关系,适合复杂场景重识别任务。
多模态物体重识别旨在利用不同模态间的互补信息检索特定目标。然而,现有方法常依赖硬性令牌过滤或简单融合策略,导致判别性线索丢失并增加背景干扰。为此,我们提出STMI框架,包含三个核心组件:(1) 分割引导特征调制(SFM)模块,利用SAM生成的掩码通过可学习注意力调制增强前景表征并抑制背景噪声;(2) 语义令牌重分配(STR)模块,采用可学习查询令牌和自适应重分配机制,在不丢弃任何令牌的前提下提取紧凑且具信息量的表示;(3) 跨模态超图交互(CHI)模块,构建跨模态统一超图以捕捉高阶语义关系。在公开基准数据集RGBNT201、RGBNT100和MSVR310上的大量实验表明,所提STMI框架在多模态重识别场景中具有优异的有效性和鲁棒性。
原文摘要 · Abstract (English)
Multi-modal object Re-Identification (ReID) aims to exploit complementary information from different modalities to retrieve specific objects. However, existing methods often rely on hard token filtering or simple fusion strategies, which can lead to the loss of discriminative cues and increased background interference. To address these challenges, we propose STMI, a novel multi-modal learning framework consisting of three key components: (1) Segmentation-Guided Feature Modulation (SFM) module leverages SAM-generated masks to enhance foreground representations and suppress background noise through learnable attention modulation; (2) Semantic Token Reallocation (STR) module employs learnable query tokens and an adaptive reallocation mechanism to extract compact and informative representations without discarding any tokens; (3) Cross-Modal Hypergraph Interaction (CHI) module constructs a unified hypergraph across modalities to capture high-order semantic relationships. Extensive experiments on public benchmarks (i.e., RGBNT201, RGBNT100, and MSVR310) demonstrate the effectiveness and robustness of our proposed STMI framework in multi-modal ReID scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。