arXiv:2604.07740cs.CVcs.AI2026-04

用文字描述提升视频行人重识别,尤其擅长复杂场景

Beyond Pedestrians: Caption-Guided CLIP Framework for High-Difficulty Video-based Person Re-Identification

  • 用文本描述和可学习标记增强特征表示
  • 在四个数据集上均超越现有方法,最高提升达12.3%
  • 适合处理服装相似、动作复杂的视频重识别任务

近年来,基于视频的行人重识别(ReID)因其能利用时空线索匹配跨非重叠摄像头的个体而受到关注。然而,当前方法在高难度场景下表现不佳,如体育和舞蹈表演中,多人穿着相似服装并进行动态动作。为此,我们提出CG-CLIP框架,通过显式文本描述与可学习标记来提升性能。该方法引入两个关键组件:基于文本的内存精炼(CMR)用于细化特定身份特征,捕捉细微差异;基于标记的特征提取(TFE)采用固定长度可学习标记的交叉注意力机制,高效聚合时空特征,降低计算开销。我们在两个标准数据集(MARS和iLIDS-VID)以及两个新构建的高难度数据集(SportsVReID和DanceVReID)上评估了该方法。实验结果表明,本方法在所有基准测试中均优于现有最先进方法,取得显著提升。

原文摘要 · Abstract (English)

In recent years, video-based person Re-Identification (ReID) has gained attention for its ability to leverage spatiotemporal cues to match individuals across non-overlapping cameras. However, current methods struggle with high-difficulty scenarios, such as sports and dance performances, where multiple individuals wear similar clothing while performing dynamic movements. To overcome these challenges, we propose CG-CLIP, a novel caption-guided CLIP framework that leverages explicit textual descriptions and learnable tokens. Our method introduces two key components: Caption-guided Memory Refinement (CMR) and Token-based Feature Extraction (TFE). CMR utilizes captions generated by Multi-modal Large Language Models (MLLMs) to refine identity-specific features, capturing fine-grained details. TFE employs a cross-attention mechanism with fixed-length learnable tokens to efficiently aggregate spatiotemporal features, reducing computational overhead. We evaluate our approach on two standard datasets (MARS and iLIDS-VID) and two newly constructed high-difficulty datasets (SportsVReID and DanceVReID). Experimental results demonstrate that our method outperforms current state-of-the-art approaches, achieving significant improvements across all benchmarks.

视频重识别文本引导多模态动态动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。