利用大模型中间层特征提升步态识别性能
BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision Models
- 提取大视觉模型多层特征融合提升识别能力
- 在多个数据集上实现超越现有方法的准确率
- 适用于跨域与同域任务,适合新研究者快速基准对比
基于大视觉模型(LVM)的步态识别已取得显著进展。然而,现有方法可能过度依赖步态先验,忽视了LVM自身多层中丰富的独特表征。本文分析发现,LVM中间层特征在不同任务中具有互补性,融合这些特征即使不依赖精心设计的步态先验也能带来显著性能提升。基于此,我们提出一个简单通用的基线方法BiggerGait。在CCPG、CAISA-B*、SUSTech1K和CCGR_MINI上的全面评估验证了BiggerGait在跨域与同域任务中的优越性,确立其作为步态表征学习的实用基准。所有模型与代码将公开发布。
原文摘要 · Abstract (English)
Large vision models (LVM) based gait recognition has achieved impressive performance. However, existing LVM-based approaches may overemphasize gait priors while neglecting the intrinsic value of LVM itself, particularly the rich, distinct representations across its multi-layers. To adequately unlock LVM's potential, this work investigates the impact of layer-wise representations on downstream recognition tasks. Our analysis reveals that LVM's intermediate layers offer complementary properties across tasks, integrating them yields an impressive improvement even without rich well-designed gait priors. Building on this insight, we propose a simple and universal baseline for LVM-based gait recognition, termed BiggerGait. Comprehensive evaluations on CCPG, CAISA-B*, SUSTech1K, and CCGR\_MINI validate the superiority of BiggerGait across both within- and cross-domain tasks, establishing it as a simple yet practical baseline for gait representation learning. All the models and code will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。