arXiv:2511.17964cs.CV2025-11AAAI被引 6

跨模态特征协作+多粒度交互,提升可见光红外视频行人重识别

X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-Identification

  • 用跨模态原型协作对齐可见光与红外特征
  • 融合短时帧间、长时跨帧及跨模态信息,增强时序建模
  • 在两个大规模数据集上超越现有方法,适合视频跨模态识别研究

大型视觉语言模型(如CLIP)在检索任务中表现优异,但其在基于视频的可见光-红外行人重识别(VVI-ReID)中的潜力尚未充分探索。主要挑战在于缩小模态差异并利用视频序列中的时空信息。为此,本文提出一种新型跨模态特征学习框架X-ReID。首先设计跨模态原型协作(CPC),对齐并整合不同模态特征,引导网络减少模态偏差;其次提出多粒度信息交互(MII),融合相邻帧的短期交互、跨帧的长期信息融合以及跨模态特征对齐,以增强时序建模并进一步缩小模态差距;最终通过整合多粒度信息,获得鲁棒的序列级表示。在两个大规模VVI-ReID基准数据集(HITSZ-VCM和BUPTCampus)上的大量实验表明,该方法优于当前最优方法。源代码已公开于https://github.com/AsuradaYuci/X-ReID。

原文摘要 · Abstract (English)

Large-scale vision-language models (e.g., CLIP) have recently achieved remarkable performance in retrieval tasks, yet their potential for Video-based Visible-Infrared Person Re-Identification (VVI-ReID) remains largely unexplored. The primary challenges are narrowing the modality gap and leveraging spatiotemporal information in video sequences. To address the above issues, in this paper, we propose a novel cross-modality feature learning framework named X-ReID for VVI-ReID. Specifically, we first propose a Cross-modality Prototype Collaboration (CPC) to align and integrate features from different modalities, guiding the network to reduce the modality discrepancy. Then, a Multi-granularity Information Interaction (MII) is designed, incorporating short-term interactions from adjacent frames, long-term cross-frame information fusion, and cross-modality feature alignment to enhance temporal modeling and further reduce modality gaps. Finally, by integrating multi-granularity information, a robust sequence-level representation is achieved. Extensive experiments on two large-scale VVI-ReID benchmarks (i.e., HITSZ-VCM and BUPTCampus) demonstrate the superiority of our method over state-of-the-art methods. The source code is released at https://github.com/AsuradaYuci/X-ReID.

视频重识别跨模态时序建模红外识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。