解决俯视角度变化下的步态识别难题,提升监控场景适应性。
CVVNet: A Cross-Vertical-View Network for Gait Recognition
- 设计多频特征提取模块,融合高低频信息应对视角变形。
- 在DroneGait和Gait3D上分别提升8.6%和2%准确率。
- 适合跨视角监控、安防系统中的步态识别任务。
步态识别可实现无接触、远距离的人体身份识别,对服装变化和非合作场景具有鲁棒性。现有方法在受控室内环境表现良好,但在俯仰角差异大的跨垂直视角场景中性能显著下降,实验显示低视角到高视角设置下准确率最高下降60%,主要由于关键解剖特征的严重形变和自遮挡。当前基于CNN和自注意力的方法因依赖单尺度卷积或简单注意力机制,难以有效整合多频率特征。为此,本文提出CVVNet(跨垂直视角网络),一种专为鲁棒跨视角步态识别设计的频率聚合架构。CVVNet采用高低频提取模块(HLFE),通过并行多尺度卷积/最大池化路径与自注意力路径分别作为高频和低频混合器,从输入轮廓中有效提取多频特征。同时引入动态门控融合(DGA)机制,自适应调节高低频特征的融合比例。结合核心的多尺度注意力门控融合(MSAGA)模块、HLFE与DGA,CVVNet能有效处理视角变化带来的形变问题,显著提升不同垂直视角下的识别鲁棒性。实验结果表明,所提方法在DroneGait上相比最优现有方法提升8.6%,在Gait3D上提升2%,达到当前最佳性能。
原文摘要 · Abstract (English)
Gait recognition enables contact-free, long-range person identification that is robust to clothing variations and non-cooperative scenarios. While existing methods perform well in controlled indoor environments, they struggle with cross-vertical view scenarios, where surveillance angles vary significantly in elevation. Our experiments show up to 60\% accuracy degradation in low-to-high vertical view settings due to severe deformations and self-occlusions of key anatomical features. Current CNN and self-attention-based methods fail to effectively handle these challenges, due to their reliance on single-scale convolutions or simplistic attention mechanisms that lack effective multi-frequency feature integration. To tackle this challenge, we propose CVVNet (Cross-Vertical-View Network), a frequency aggregation architecture specifically designed for robust cross-vertical-view gait recognition. CVVNet employs a High-Low Frequency Extraction module (HLFE) that adopts parallel multi-scale convolution/max-pooling path and self-attention path as high- and low-frequency mixers for effective multi-frequency feature extraction from input silhouettes. We also introduce the Dynamic Gated Aggregation (DGA) mechanism to adaptively adjust the fusion ratio of high- and low-frequency features. The integration of our core Multi-Scale Attention Gated Aggregation (MSAGA) module, HLFE and DGA enables CVVNet to effectively handle distortions from view changes, significantly improving the recognition robustness across different vertical views. Experimental results show that our CVVNet achieves state-of-the-art performance, with $8.6\%$ improvement on DroneGait and $2\%$ on Gait3D compared with the best existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。