用人体骨骼引导视频特征学习,提升可见光与红外行人重识别精度
Skeleton-Guided Spatial-Temporal Feature Learning for Video-Based Visible-Infrared Person Re-Identification
- 基于骨骼信息在帧级和序列级优化时空特征
- 在多个基准数据集上超越现有最先进方法
- 特别适合处理低质量、遮挡严重的红外视频
基于视频的可见光-红外行人重识别(VVI-ReID)因模态间特征差异大而面临挑战。视频中的时空信息至关重要,但其准确性常受视频质量低、遮挡等问题影响。现有方法主要关注减少模态差异,却较少关注提升时空特征,尤其对红外视频。为此,我们提出一种新型骨架引导时空特征学习(STAR)方法。利用对图像质量差、遮挡等鲁棒的骨架信息,增强双模态视频的时空特征精度。STAR采用两级骨架引导策略:帧级通过结构化骨架信息细化单帧视觉特征;序列级设计基于骨架关键点图的特征聚合机制,学习不同身体部位对时空特征的贡献,进一步提升全局特征精度。在多个基准数据集上的实验表明,STAR优于现有最先进方法。代码即将开源。
原文摘要 · Abstract (English)
Video-based visible-infrared person re-identification (VVI-ReID) is challenging due to significant modality feature discrepancies. Spatial-temporal information in videos is crucial, but the accuracy of spatial-temporal information is often influenced by issues like low quality and occlusions in videos. Existing methods mainly focus on reducing modality differences, but pay limited attention to improving spatial-temporal features, particularly for infrared videos. To address this, we propose a novel Skeleton-guided spatial-Temporal feAture leaRning (STAR) method for VVI-ReID. By using skeleton information, which is robust to issues such as poor image quality and occlusions, STAR improves the accuracy of spatial-temporal features in videos of both modalities. Specifically, STAR employs two levels of skeleton-guided strategies: frame level and sequence level. At the frame level, the robust structured skeleton information is used to refine the visual features of individual frames. At the sequence level, we design a feature aggregation mechanism based on skeleton key points graph, which learns the contribution of different body parts to spatial-temporal features, further enhancing the accuracy of global features. Experiments on benchmark datasets demonstrate that STAR outperforms state-of-the-art methods. Code will be open source soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。