用DINOv2学习步态特征,提升可见光与红外视频行人重识别性能
DINOv2 Driven Gait Representation Learning for Video-Based Visible-Infrared Person Re-identification
- 利用DINOv2语义先验增强步态轮廓,联合优化身份识别任务
- 通过多粒度双向交互,融合步态与外观特征,提升全局表征能力
- 适合跨模态视频行人重识别研究者,尤其关注动态特征建模的场景
基于视频的可见光-红外行人重识别(VVI-ReID)旨在从视频序列中跨模态检索同一行人。现有方法多聚焦于模态不变的视觉特征,却忽视了兼具模态不变性与丰富时序动态性的步态特征,限制了对时空一致性的建模能力。为此,我们提出基于DINOv2的步态表征学习框架(DinoGRL),利用DINOv2丰富的视觉先验,学习与外观互补的步态特征,实现鲁棒的序列级表征。具体地,设计语义感知轮廓与步态学习模型(SASGL),通过DINOv2提供的通用语义先验生成并增强轮廓表示,并联合优化身份识别目标,实现语义丰富且任务自适应的步态特征学习。此外,构建渐进式双向多粒度增强模块(PBMGE),在多空间粒度上实现步态与外观流的双向交互,充分挖掘二者互补性,以局部细节增强全局表征,生成高度判别性特征。在HITSZ-VCM和BUPT数据集上的大量实验表明,该方法显著优于现有最先进方法。
原文摘要 · Abstract (English)
Video-based Visible-Infrared person re-identification (VVI-ReID) aims to retrieve the same pedestrian across visible and infrared modalities from video sequences. Existing methods tend to exploit modality-invariant visual features but largely overlook gait features, which are not only modality-invariant but also rich in temporal dynamics, thus limiting their ability to model the spatiotemporal consistency essential for cross-modal video matching. To address these challenges, we propose a DINOv2-Driven Gait Representation Learning (DinoGRL) framework that leverages the rich visual priors of DINOv2 to learn gait features complementary to appearance cues, facilitating robust sequence-level representations for cross-modal retrieval. Specifically, we introduce a Semantic-Aware Silhouette and Gait Learning (SASGL) model, which generates and enhances silhouette representations with general-purpose semantic priors from DINOv2 and jointly optimizes them with the ReID objective to achieve semantically enriched and task-adaptive gait feature learning. Furthermore, we develop a Progressive Bidirectional Multi-Granularity Enhancement (PBMGE) module, which progressively refines feature representations by enabling bidirectional interactions between gait and appearance streams across multiple spatial granularities, fully leveraging their complementarity to enhance global representations with rich local details and produce highly discriminative features. Extensive experiments on HITSZ-VCM and BUPT datasets demonstrate the superiority of our approach, significantly outperforming existing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。