arXiv:2409.09391cs.CV2024-09被引 7

融合姿态与全局关系的图网络,提升监控视频中行人重识别准确率

Tran-GCN: A Transformer-Enhanced Graph Convolutional Network for Person Re-Identification in Monitoring Videos

  • 用姿态估计和Transformer捕捉行人局部特征间的全局依赖
  • 在Market-1501等三个数据集上显著提升识别准确率
  • 适合关注跨摄像头行人识别的安防与视觉系统研究者

行人重识别(Re-ID)在计算机视觉中日益重要,实现跨摄像头行人识别。尽管深度学习为该领域提供了坚实基础,但现有方法常忽略局部特征间的潜在关系,难以应对姿态变化和局部遮挡问题。为此,我们提出一种增强型图卷积网络Tran-GCN,包含四个关键组件:(1) 姿态估计学习分支用于提取行人的关键点信息;(2) Transformer学习分支建模细粒度语义局部特征间的全局依赖;(3) 卷积学习分支采用基础ResNet架构提取局部细节特征;(4) 图卷积模块(GCM)融合局部、全局特征与身体结构信息,实现更有效的识别。在Market-1501、DukeMTMC-ReID和MSMT17三个数据集上的定量与定性实验表明,Tran-GCN能更准确捕获监控视频中的判别性特征,显著提升识别精度。

原文摘要 · Abstract (English)

Person Re-Identification (Re-ID) has gained popularity in computer vision, enabling cross-camera pedestrian recognition. Although the development of deep learning has provided a robust technical foundation for person Re-ID research, most existing person Re-ID methods overlook the potential relationships among local person features, failing to adequately address the impact of pedestrian pose variations and local body parts occlusion. Therefore, we propose a Transformer-enhanced Graph Convolutional Network (Tran-GCN) model to improve Person Re-Identification performance in monitoring videos. The model comprises four key components: (1) A Pose Estimation Learning branch is utilized to estimate pedestrian pose information and inherent skeletal structure data, extracting pedestrian key point information; (2) A Transformer learning branch learns the global dependencies between fine-grained and semantically meaningful local person features; (3) A Convolution learning branch uses the basic ResNet architecture to extract the person's fine-grained local features; (4) A Graph Convolutional Module (GCM) integrates local feature information, global feature information, and body information for more effective person identification after fusion. Quantitative and qualitative analysis experiments conducted on three different datasets (Market-1501, DukeMTMC-ReID, and MSMT17) demonstrate that the Tran-GCN model can more accurately capture discriminative person features in monitoring videos, significantly improving identification accuracy.

行人重识别图神经网络Transformer姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。