arXiv:2604.27353cs.CV2026-04

通过多分支融合提升步态识别精度,尤其在变装和视角变化下表现优异。

Gait Recognition via Deep Residual Networks and Multi-Branch Feature Fusion

  • 用多分支结构分别提取体型、步速和骨骼运动特征
  • 在CASIA-B数据集上正常行走时达到94.52%的Rank-1准确率
  • 适合安防监控中远距离、非接触式身份识别场景

步态识别作为监视与安全应用中的重要生物特征,具有非侵入性、抗伪装和远距离识别等优势。然而现有方法难以全面捕捉人体运动中的丰富生物特征,尤其在视角变化、着装改变和携带物品等干扰条件下表现不佳。本文提出一种高精度步态识别框架,基于深度残差学习的多分支架构,深度融合步态动态与体态特征。首先采用高分辨率网络(HRNet)进行鲁棒的骨骼关键点估计,保留低分辨率输入下的细粒度空间信息。随后从姿态序列构建三个互补特征分支:体型比例、步态速度和骨骼运动。利用50层残差网络(ResNet-50)主干提取分层且具有判别力的特征表示。为有效融合异构特征流,设计了受通道注意力启发的多分支特征融合(MFF)模块,通过可学习激活参数动态分配各分支贡献权重。在跨视角多条件的CASIA-B基准测试中,本方法在正常行走条件下达到94.52%的Rank-1准确率,是当前基于骨架的方法中在穿外套条件下的最佳性能。

原文摘要 · Abstract (English)

Gait recognition has emerged as a compelling biometric modality for surveillance and security applications, offering inherent advantages such as non-intrusiveness, resistance to disguise, and long-range identification capability. However, prevailing approaches struggle to comprehensively capture and exploit the rich biometric cues embedded in human locomotion, particularly under covariate interference including viewpoint variation, clothing change, and carrying conditions. In this paper, we present a high-precision gait recognition framework that deeply extracts and synergistically fuses gait dynamics with body shape characteristics through a multi-branch architecture grounded in deep residual learning. Specifically, we first employ the High-Resolution Network (HRNet) to perform robust skeletal keypoint estimation, preserving fine-grained spatial information even under low-resolution inputs. We then construct three complementary feature branches -- body proportion, gait velocity, and skeletal motion -- from the extracted pose sequences. A 50-layer Residual Network (ResNet-50) backbone is leveraged within a deep feature extraction module to capture hierarchically rich and discriminative representations. To effectively integrate heterogeneous feature streams, we design a Multi-Branch Feature Fusion (MFF) module inspired by channel-wise attention mechanisms, which dynamically allocates contribution weights across branches through learned activation parameters. Extensive experiments on the cross-view multi-condition CASIA-B benchmark demonstrate that our method achieves a Rank-1 accuracy of 94.52\% under normal walking, with the best recognition performance among skeleton-based methods for the coat-wearing condition.

步态识别多分支融合残差网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。