通过配对匹配建模可见光与红外图像的双向交互,提升跨模态行人重识别效果。
BIT: Matching-based Bi-directional Interaction Transformation Network for Visible-Infrared Person Re-Identification
- 基于配对匹配机制,显式建模可见光与红外图像间的双向交互。
- 在多个基准上达到最新性能,尤其在红外样本稀缺时表现更优。
- 适合关注跨模态匹配与小样本分布偏移问题的研究者。
可见光-红外行人重识别(VI-ReID)因两种模态间存在显著差异而具有挑战性。现有方法虽尝试在共享嵌入空间中学习模态不变特征以弥合差距,但常忽略模态间复杂的隐含关联。这一缺陷在分布偏移场景下尤为严重,此时红外样本数量远少于可见光样本。为此,本文提出双向交互变换网络(BIT),摒弃刚性特征对齐,采用基于配对匹配的策略,显式建模可见光与红外图像对之间的交互。BIT采用编码器-解码器结构:编码器提取初步特征表示,解码器执行双向特征融合与查询感知评分,强化跨模态对应关系。据我们所知,BIT是首个在VI-ReID中引入此类配对驱动交互的方法。大量实验表明,该方法在多个基准上达到最优性能,充分验证了其有效性。
原文摘要 · Abstract (English)
Visible-Infrared Person Re-Identification (VI-ReID) is a challenging retrieval task due to the substantial modality gap between visible and infrared images. While existing methods attempt to bridge this gap by learning modality-invariant features within a shared embedding space, they often overlook the complex and implicit correlations between modalities. This limitation becomes more severe under distribution shifts, where infrared samples are often far fewer than visible ones. To address these challenges, we propose a novel network termed Bi-directional Interaction Transformation (BIT). Instead of relying on rigid feature alignment, BIT adopts a matching-based strategy that explicitly models the interaction between visible and infrared image pairs. Specifically, BIT employs an encoder-decoder architecture where the encoder extracts preliminary feature representations, and the decoder performs bi-directional feature integration and query aware scoring to enhance cross-modality correspondence. To our best knowledge, BIT is the first to introduce such pairwise matching-driven interaction in VI-ReID. Extensive experiments on several benchmarks demonstrate that our BIT achieves state-of-the-art performance, highlighting its effectiveness in the VI-ReID task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。