融合2D与3D姿态信息,实现远距离下的人体步态识别与属性分析。
Combo-Gait: Unified Transformer Framework for Multi-Modal Gait Recognition and Attribute Analysis
- 用统一Transformer融合2D轮廓和3D SMPL特征
- 在1公里远、50度俯仰角下仍保持高识别准确率
- 同时预测年龄、体重指数和性别,适合实际监控场景
步态识别是远距离人体识别的重要生物特征,尤其在低分辨率或非约束环境下。现有方法多依赖单一模态(如2D轮廓或3D网格),难以全面捕捉行走的几何与动态复杂性。本文提出一种多模态、多任务框架,结合2D时序轮廓与3D SMPL特征,实现鲁棒步态分析。除身份识别外,还引入多任务学习策略,联合完成年龄、身体质量指数(BMI)和性别估计。采用统一Transformer有效融合多模态特征,同时保留身份判别线索。在大规模BRIAR数据集上进行实验,该数据集涵盖最长1公里距离和最高50°俯仰角等挑战性条件,结果表明本方法在步态识别上超越现有最佳模型,并实现精准的人体属性估计。验证了多模态与多任务学习在真实场景中推进步态理解的潜力。
原文摘要 · Abstract (English)
Gait recognition is an important biometric for human identification at a distance, particularly under low-resolution or unconstrained environments. Current works typically focus on either 2D representations (e.g., silhouettes and skeletons) or 3D representations (e.g., meshes and SMPLs), but relying on a single modality often fails to capture the full geometric and dynamic complexity of human walking patterns. In this paper, we propose a multi-modal and multi-task framework that combines 2D temporal silhouettes with 3D SMPL features for robust gait analysis. Beyond identification, we introduce a multitask learning strategy that jointly performs gait recognition and human attribute estimation, including age, body mass index (BMI), and gender. A unified transformer is employed to effectively fuse multi-modal gait features and better learn attribute-related representations, while preserving discriminative identity cues. Extensive experiments on the large-scale BRIAR datasets, collected under challenging conditions such as long-range distances (up to 1 km) and extreme pitch angles (up to 50°), demonstrate that our approach outperforms state-of-the-art methods in gait recognition and provides accurate human attribute estimation. These results highlight the promise of multi-modal and multitask learning for advancing gait-based human understanding in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。