用语言描述+运动感知,提升步态识别泛化能力
Language-Guided and Motion-Aware Gait Representation for Generalizable Recognition
- 引入自然语言作为语义先验,定位关键运动区域
- 在三个数据集上准确率超97%,显著优于现有方法
- 适合需要跨环境、抗干扰步态识别的场景
步态识别是计算机视觉中一种有前景的技术,广泛应用于远程人体识别。然而,现有方法通常依赖复杂架构直接从图像提取特征,并通过池化获得序列级表示,易受静态噪声(如衣物)干扰,难以有效捕捉动态运动区域(如手臂和腿部)。这一瓶颈在类内差异显著时尤为突出,同一人在不同环境下的步态特征在特征空间中相距甚远。为此,我们提出首个引入自然语言描述作为显式语义先验的步态识别框架——LMGait。通过设计与步态相关的语言线索,捕捉步态序列中的关键运动特征。为增强跨模态对齐,提出运动感知模块(MAM),自适应调整语义信息以更好对齐视觉表征。此外,引入运动时序捕获模块(MTCM),增强特征判别力并提升运动追踪能力。在多个数据集上的实验表明,该模型在CCPG、SUSTech1K和CASIAB上分别达到88.5%、97.1%和97.5%的准确率,达到当前最优性能。
原文摘要 · Abstract (English)
Gait recognition is emerging as a promising technology and an innovative field within computer vision, with a wide range of applications in remote human identification. However, existing methods typically rely on complex architectures to directly extract features from images and apply pooling operations to obtain sequence-level representations. Such designs often lead to overfitting on static noise (e.g., clothing), while failing to effectively capture dynamic motion regions, such as the arms and legs. This bottleneck is particularly challenging in the presence of intra-class variation, where gait features of the same individual under different environmental conditions are significantly distant in the feature space. To address the above challenges, we present a Languageguided and Motion-aware gait recognition framework, named LMGait. To the best of our knowledge, LMGait is the first method to introduce natural language descriptions as explicit semantic priors into the gait recognition task. In particular, we utilize designed gait-related language cues to capture key motion features in gait sequences. To improve cross-modal alignment, we propose the Motion Awareness Module (MAM), which refines the language features by adaptively adjusting various levels of semantic information to ensure better alignment with the visual representations. Furthermore, we introduce the Motion Temporal Capture Module (MTCM) to enhance the discriminative capability of gait features and improve the model's motion tracking ability. We conducted extensive experiments across multiple datasets, and the results demonstrate the significant advantages of our proposed network. Specifically, our model achieved accuracies of 88.5%, 97.1%, and 97.5% on the CCPG, SUSTech1K, and CASIAB datasets, respectively, achieving state-of-the-art performance. Homepage: https://dingwu1021.github.io/LMGait/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。