arXiv:2508.21363cs.CV2025-08中稿 · IEEE TRANSACTIONS …被引 3

用分层时间剪枝提升扩散模型3D人体姿态估计效率

Efficient Diffusion-Based 3D Human Pose Estimation with Hierarchical Temporal Pruning

  • 分三阶段动态剪枝:帧级相关性识别、稀疏注意力聚焦、语义聚类精简
  • 推理计算量降低56.8%,速度提升81.1%,性能达当前最优
  • 适合追求高效高精度3D姿态估计的视觉与动作分析研究者

扩散模型在生成高保真3D人体姿态方面表现优异,但其迭代特性与多假设需求导致计算开销巨大。本文提出一种基于分层时间剪枝(HTP)的高效扩散3D人体姿态估计框架,通过帧级与语义级双重动态剪枝,在保留关键运动动态的同时减少冗余姿态标记。HTP分三步执行:(1) 通过自适应时序图构建分析帧间运动相关性,实现时序相关增强剪枝(TCEP);(2) 利用帧级稀疏性,采用稀疏聚焦时序多头自注意力(SFT MHSA),仅关注运动相关标记以降低注意力计算量;(3) 基于聚类的掩码引导姿态标记剪枝(MGPTP)进行细粒度语义剪枝,保留最具信息量的姿态标记。在Human3.6M与MPI-INF-3DHP数据集上的实验表明,相比先前扩散方法,HTP将训练MACs降低38.5%,推理MACs降低56.8%,平均推理速度提升81.1%,同时达到顶尖性能。

原文摘要 · Abstract (English)

Diffusion models have demonstrated strong capabilities in generating high-fidelity 3D human poses, yet their iterative nature and multi-hypothesis requirements incur substantial computational cost. In this paper, we propose an Efficient Diffusion-Based 3D Human Pose Estimation framework with a Hierarchical Temporal Pruning (HTP) strategy, which dynamically prunes redundant pose tokens across both frame and semantic levels while preserving critical motion dynamics. HTP operates in a staged, top-down manner: (1) Temporal Correlation-Enhanced Pruning (TCEP) identifies essential frames by analyzing inter-frame motion correlations through adaptive temporal graph construction; (2) Sparse-Focused Temporal MHSA (SFT MHSA) leverages the resulting frame-level sparsity to reduce attention computation, focusing on motion-relevant tokens; and (3) Mask-Guided Pose Token Pruner (MGPTP) performs fine-grained semantic pruning via clustering, retaining only the most informative pose tokens. Experiments on Human3.6M and MPI-INF-3DHP show that HTP reduces training MACs by 38.5\%, inference MACs by 56.8\%, and improves inference speed by an average of 81.1\% compared to prior diffusion-based methods, while achieving state-of-the-art performance.

3D姿态估计扩散模型效率优化时序剪枝

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。