通过融合多几何特征提升动作计数精度,解决视角变化导致的误检问题。
GMFL-Net: A Global Multi-geometric Feature Learning Network for Repetitive Action Counting
- 融合多几何特征并学习语义相似性,增强表征能力
- 全局建模点与通道间依赖关系,生成综合特征表示
- 新数据集含细粒度标注,适合复杂动作计数研究
随着深度学习的发展,重复性动作计数逐渐受到关注。基于人体姿态估计网络提取姿态关键点的方法已被证明在姿态级别上有效。然而,现有方法因单个坐标不稳定,在相机视角变化下难以准确识别显著姿态,且在异常到实际动作的过渡阶段易发生误检。为此,本文提出一种简单高效的全局多几何特征学习网络(GMFL-Net)。设计MIA-Module,通过融合多几何特征并学习其语义相似性来提升信息表征;同时设计GBFL-Module,从全局角度增强点间与通道间依赖关系,并结合MIA-Module生成的丰富局部信息,合成全面且最具代表性的全局特征表示。此外,针对现有数据集不足的问题,构建新数据集Countix-Fitness-pose,包含不同周期长度与异常情况,测试集持续时间更长,并在姿态级别进行细粒度标注,新增深蹲与跳绳推举两个动作类别。在RepCount-pose、UCFRep-pose和Countix-Fitness-pose三个挑战性基准上的实验表明,所提GMFL-Net达到当前最优性能。
原文摘要 · Abstract (English)
With the continuous development of deep learning, the field of repetitive action counting is gradually gaining notice from many researchers. Extraction of pose keypoints using human pose estimation networks is proven to be an effective pose-level method. However, existing pose-level methods suffer from the shortcomings that the single coordinate is not stable enough to handle action distortions due to changes in camera viewpoints, thus failing to accurately identify salient poses, and is vulnerable to misdetection during the transition from the exception to the actual action. To overcome these problems, we propose a simple but efficient Global Multi-geometric Feature Learning Network (GMFL-Net). Specifically, we design a MIA-Module that aims to improve information representation by fusing multi-geometric features, and learning the semantic similarity among the input multi-geometric features. Then, to improve the feature representation from a global perspective, we also design a GBFL-Module that enhances the inter-dependencies between point-wise and channel-wise elements and combines them with the rich local information generated by the MIA-Module to synthesise a comprehensive and most representative global feature representation. In addition, considering the insufficient existing dataset, we collect a new dataset called Countix-Fitness-pose (https://github.com/Wantong66/Countix-Fitness) which contains different cycle lengths and exceptions, a test set with longer duration, and annotate it with fine-grained annotations at the pose-level. We also add two new action classes, namely lunge and rope push-down. Finally, extensive experiments on the challenging RepCount-pose, UCFRep-pose, and Countix-Fitness-pose benchmarks show that our proposed GMFL-Net achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。