解决远距离步态识别难题,融合摄像头与激光雷达数据提升识别鲁棒性。
Walking Further: Semantic-aware Multimodal Gait Recognition Under Long-Range Conditions

- 用语义引导融合框架对齐图像与点云特征,提升跨模态一致性。
- 在长距离多场景下达到92.3%识别准确率,显著优于现有方法。
- 适合安防、远程监控等真实户外环境中的身份识别应用。
步态识别是一种新兴的生物特征技术,可实现非侵入式且难伪造的人体识别。然而,现有方法大多局限于短距离、单模态场景,在真实环境下远距离和跨距离情形下泛化能力差。为此,我们提出首个面向远距离步态识别的激光雷达-摄像头多模态基准数据集LRGait,同时设计端到端的EMGaitNet框架。为弥合RGB图像与点云之间的模态差异,引入语义引导融合流程:基于CLIP的语义挖掘(SeMi)模块提取人体部位感知的语义线索,并通过语义引导对齐(SGA)模块在统一嵌入空间中对齐2D与3D特征;对称交叉注意力融合(SCAF)模块分层整合视觉轮廓与3D几何特征,时空(ST)模块捕捉全局步态动态。在多个步态数据集上的大量实验验证了方法的有效性。
原文摘要 · Abstract (English)
Gait recognition is an emerging biometric technology that enables non-intrusive and hard-to-spoof human identification. However, most existing methods are confined to short-range, unimodal settings and fail to generalize to long-range and cross-distance scenarios under real-world conditions. To address this gap, we present \textbf{LRGait}, the first LiDAR-Camera multimodal benchmark designed for robust long-range gait recognition across diverse outdoor distances and environments. We further propose \textbf{EMGaitNet}, an end-to-end framework tailored for long-range multimodal gait recognition. To bridge the modality gap between RGB images and point clouds, we introduce a semantic-guided fusion pipeline. A CLIP-based Semantic Mining (SeMi) module first extracts human body-part-aware semantic cues, which are then employed to align 2D and 3D features via a Semantic-Guided Alignment (SGA) module within a unified embedding space. A Symmetric Cross-Attention Fusion (SCAF) module hierarchically integrates visual contours and 3D geometric features, and a Spatio-Temporal (ST) module captures global gait dynamics. Extensive experiments on various gait datasets validate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。