构建多模态步态识别基准,支持跨模态统一识别
MMGait: Towards Multi-Modal Gait Recognition

- 整合五类传感器数据,建立包含12种模态的多模态步态数据库
- 在725人、33.4万条序列上验证,实现跨模态识别性能提升
- 提出OmniGait模型,统一单/跨/多模态识别任务,适合实际部署
步态识别作为远距离非接触式身份识别技术备受关注。现有方法主要依赖RGB图像,难以应对真实场景中多模态协同与跨模态检索需求。为此,我们提出MMGait,一个涵盖五个异构传感器(可见光相机、深度相机、红外相机、激光雷达、4D雷达)的综合性多模态步态基准。该数据集包含12种模态、334,060条序列,来自725名受试者,支持对几何、光照和运动特征的系统研究。基于此,我们在单模态、跨模态和多模态三种范式下进行广泛评估,分析各模态的鲁棒性与互补性。此外,我们引入全新任务——全模态步态识别(Omni Multi-Modal Gait Recognition),旨在统一上述三类识别模式。我们还提出简单而有效的基线模型OmniGait,通过学习跨模态共享嵌入空间,在多种设置下取得优异表现。相关数据集、代码与预训练模型已开源。
原文摘要 · Abstract (English)
Gait recognition has emerged as a powerful biometric technique for identifying individuals at a distance without requiring user cooperation. Most existing methods focus primarily on RGB-derived modalities, which fall short in real-world scenarios requiring multi-modal collaboration and cross-modal retrieval. To overcome these challenges, we present MMGait, a comprehensive multi-modal gait benchmark integrating data from five heterogeneous sensors, including an RGB camera, a depth camera, an infrared camera, a LiDAR scanner, and a 4D Radar system. MMGait contains twelve modalities and 334,060 sequences from 725 subjects, enabling systematic exploration across geometric, photometric, and motion domains. Based on MMGait, we conduct extensive evaluations on single-modal, cross-modal, and multi-modal paradigms to analyze modality robustness and complementarity. Furthermore, we introduce a new task, Omni Multi-Modal Gait Recognition, which aims to unify the above three gait recognition paradigms within a single model. We also propose a simple yet powerful baseline, OmniGait, which learns a shared embedding space across diverse modalities and achieves promising recognition performance. The MMGait benchmark, codebase, and pretrained checkpoints are publicly available at https://github.com/BNU-IVC/MMGait.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。