arXiv:2501.15073cs.CV2025-01AAAI被引 5

用稀疏标注视频提升人体姿态估计,显著降低标注成本。

SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos

  • 引入动态感知掩码捕捉长期运动上下文
  • 在三大数据集上达到新基准,仅需26.7%标注数据
  • 适合低标注成本、长时序动作分析场景

视频中的人体姿态估计仍面临挑战,主要源于大规模数据集依赖人工标注,成本高昂。现有方法常难以捕捉长时序依赖关系,且忽略姿态热图与视觉特征间的互补性。为此,本文提出STDPose框架,通过学习稀疏标注视频中的时空动态来提升姿态估计性能。该框架包含两项关键创新:1)动态感知掩码,用于捕捉长时运动上下文,实现对姿态变化的精细理解;2)时空表示编码与聚合系统,有效建模时空关系,增强姿态估计的准确性和鲁棒性。STDPose在三个大规模数据集上建立了视频姿态传播(即从标注帧向未标注帧传播姿态标签)和姿态估计任务的新基准。此外,利用姿态传播生成的伪标签,仅使用26.7%的真实标注数据即可达到具有竞争力的性能。

原文摘要 · Abstract (English)

Human pose estimation in videos remains a challenge, largely due to the reliance on extensive manual annotation of large datasets, which is expensive and labor-intensive. Furthermore, existing approaches often struggle to capture long-range temporal dependencies and overlook the complementary relationship between temporal pose heatmaps and visual features. To address these limitations, we introduce STDPose, a novel framework that enhances human pose estimation by learning spatiotemporal dynamics in sparsely-labeled videos. STDPose incorporates two key innovations: 1) A novel Dynamic-Aware Mask to capture long-range motion context, allowing for a nuanced understanding of pose changes. 2) A system for encoding and aggregating spatiotemporal representations and motion dynamics to effectively model spatiotemporal relationships, improving the accuracy and robustness of pose estimation. STDPose establishes a new performance benchmark for both video pose propagation (i.e., propagating pose annotations from labeled frames to unlabeled frames) and pose estimation tasks, across three large-scale evaluation datasets. Additionally, utilizing pseudo-labels generated by pose propagation, STDPose achieves competitive performance with only 26.7% labeled data.

姿态估计稀疏标注时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。