用骨骼序列替代文本,让视频行人重识别更精准。
Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification
- 以骨骼序列作为核心模态,实现运动感知的预训练。
- 在MARS、LS-VID等数据集上达到新最好效果,超越现有方法。
- 适合关注动作特征建模的视频重识别研究者使用。
多模态预训练革新了视觉理解,但在基于视频的行人重识别(ReID)中影响仍有限。现有方法依赖视频-文本对,存在两大缺陷:缺乏真实多模态预训练,且文本无法捕捉精细的时间运动信息——这对区分视频中身份至关重要。本文首次提出基于骨骼序列的预训练框架。提出对比骨骼-图像预训练(CSIP-ReID),采用两阶段设计:第一阶段通过对比学习对齐骨骼与视觉特征;第二阶段引入动态原型融合更新器(PFU),融合运动与外观线索。同时提出骨骼引导时间建模(SGTM)模块,从骨骼数据中提取时间特征并融入视觉表示。大量实验表明,该方法在标准视频ReID基准(MARS、LS-VID、iLIDS-VID)上取得新最优结果,并在仅用骨骼的ReID任务(BIWI、IAS)中显著优于以往方法。本工作开创了无需标注、运动感知的预训练范式,为多模态表征学习开辟新方向。
原文摘要 · Abstract (English)
Multimodal pretraining has revolutionized visual understanding, but its impact on video-based person re-identification (ReID) remains underexplored. Existing approaches often rely on video-text pairs, yet suffer from two fundamental limitations: (1) lack of genuine multimodal pretraining, and (2) text poorly captures fine-grained temporal motion-an essential cue for distinguishing identities in video. In this work, we take a bold departure from text-based paradigms by introducing the first skeleton-driven pretraining framework for ReID. To achieve this, we propose Contrastive Skeleton-Image Pretraining for ReID (CSIP-ReID), a novel two-stage method that leverages skeleton sequences as a spatiotemporally informative modality aligned with video frames. In the first stage, we employ contrastive learning to align skeleton and visual features at sequence level. In the second stage, we introduce a dynamic Prototype Fusion Updater (PFU) to refine multimodal identity prototypes, fusing motion and appearance cues. Moreover, we propose a Skeleton Guided Temporal Modeling (SGTM) module that distills temporal cues from skeleton data and integrates them into visual features. Extensive experiments demonstrate that CSIP-ReID achieves new state-of-the-art results on standard video ReID benchmarks (MARS, LS-VID, iLIDS-VID). Moreover, it exhibits strong generalization to skeleton-only ReID tasks (BIWI, IAS), significantly outperforming previous methods. CSIP-ReID pioneers an annotation-free and motion-aware pretraining paradigm for ReID, opening a new frontier in multimodal representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。