arXiv:2601.12062cs.CV2026-01被引 1

用语言提示提升可见光与红外视频行人重识别的跨模态一致性

Learning Language-Driven Sequence-Level Modal-Invariant Representations for Video-Based Visible-Infrared Person Re-Identification

  • 基于CLIP构建轻量时空特征模块,高效建模视频序列
  • 通过语义扩散和双向交叉注意力增强跨模态交互,减少模态差异
  • 引入双模态损失,提升模型对未知类别的泛化能力

视频级可见光-红外行人重识别(VVI-ReID)的核心在于学习跨模态的序列级不变表示。现有方法多依赖CLIP生成的共享语言提示来引导学习,虽性能优异,但仍存在时空建模效率低、跨模态交互不足、缺乏显式模态级监督等问题。为此,提出语言驱动的序列级模态不变表示学习(LSMRL)方法,包含时空特征学习(STFL)、语义扩散(SD)和跨模态交互(CMI)模块。STFL模块在CLIP基础上最小改动实现高效时空建模;SD模块将共享语言提示扩散至可见光与红外特征,建立初步模态一致性;CMI模块利用双向跨模态自注意力消除残余模态差异,优化不变表示。此外,引入两种模态级损失,明确增强特征判别性与未见类别泛化能力。大规模VVI-ReID数据集上的实验表明,LSMRL显著优于当前最优方法。

原文摘要 · Abstract (English)

The core of video-based visible-infrared person re-identification (VVI-ReID) lies in learning sequence-level modal-invariant representations across different modalities. Recent research tends to use modality-shared language prompts generated by CLIP to guide the learning of modal-invariant representations. Despite achieving optimal performance, such methods still face limitations in efficient spatial-temporal modeling, sufficient cross-modal interaction, and explicit modality-level loss guidance. To address these issues, we propose the language-driven sequence-level modal-invariant representation learning (LSMRL) method, which includes spatial-temporal feature learning (STFL) module, semantic diffusion (SD) module and cross-modal interaction (CMI) module. To enable parameter- and computation-efficient spatial-temporal modeling, the STFL module is built upon CLIP with minimal modifications. To achieve sufficient cross-modal interaction and enhance the learning of modal-invariant features, the SD module is proposed to diffuse modality-shared language prompts into visible and infrared features to establish preliminary modal consistency. The CMI module is further developed to leverage bidirectional cross-modal self-attention to eliminate residual modality gaps and refine modal-invariant representations. To explicitly enhance the learning of modal-invariant representations, two modality-level losses are introduced to improve the features' discriminative ability and their generalization to unseen categories. Extensive experiments on large-scale VVI-ReID datasets demonstrate the superiority of LSMRL over AOTA methods.

行人重识别跨模态视频理解语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。