arXiv:2503.05365cs.CV2025-03

通过多粒度特征剪枝,提升视频人体姿态估计的速度与精度。

Multi-Grained Feature Pruning for Video-Based Human Pose Estimation

  • 设计多尺度分辨率框架,动态识别关键语义特征
  • 在PoseTrack2017上达87.4 mAP,推理速度提升93.8%
  • 适合需要高效高精度姿态估计的应用场景

人体姿态估计在动作识别和运动捕捉中有广泛应用,近年来取得显著进展。然而,当前基于Transformer的视频姿态估计方法常因冗余时间信息处理和细粒度感知不足而受限,主要因仅处理低分辨率特征。为此,我们提出一种新型多尺度分辨率框架,以不同粒度编码时空表征并实现细粒度感知补偿。同时,采用密度峰值聚类动态识别并优先处理携带重要语义信息的特征令牌,有效剪除冗余特征,尤其针对多帧特征,从而在不损失语义丰富性前提下优化计算效率。实验表明,该方法在三个大规模数据集上均刷新性能与效率基准,在PoseTrack2017上达到87.4 mAP,推理速度相较基线提升93.8%。

原文摘要 · Abstract (English)

Human pose estimation, with its broad applications in action recognition and motion capture, has experienced significant advancements. However, current Transformer-based methods for video pose estimation often face challenges in managing redundant temporal information and achieving fine-grained perception because they only focus on processing low-resolution features. To address these challenges, we propose a novel multi-scale resolution framework that encodes spatio-temporal representations at varying granularities and executes fine-grained perception compensation. Furthermore, we employ a density peaks clustering method to dynamically identify and prioritize tokens that offer important semantic information. This strategy effectively prunes redundant feature tokens, especially those arising from multi-frame features, thereby optimizing computational efficiency without sacrificing semantic richness. Empirically, it sets new benchmarks for both performance and efficiency on three large-scale datasets. Our method achieves a 93.8% improvement in inference speed compared to the baseline, while also enhancing pose estimation accuracy, reaching 87.4 mAP on the PoseTrack2017 dataset.

姿态估计视频分析特征剪枝Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。