基于运动学设计骨骼点拓扑关系,提升动作识别精度。
HFGCN:Hypergraph Fusion Graph Convolutional Networks for Skeleton-Based Action Recognition
- 结合运动学分析骨骼点与身体部位的拓扑关系
- 在三个数据集上超越现有最先进方法,准确率显著提升
- 适合关注骨骼动作识别与图神经网络融合的研究者
近年来,动作识别因在视频理解中的关键作用而受到广泛关注并得到广泛应用。现有研究多通过深度学习方法提升性能,却忽视了骨骼点之间的拓扑建模。尽管部分工作采用数据驱动方式构建骨骼点拓扑,但未考虑骨骼点的运动学特性。为此,本文借鉴运动学理论,提出基于身体部位与距躯干中心距离的骨骼点拓扑关系分类方法。为融合该拓扑信息用于动作识别,提出一种新型超图融合图卷积网络(HFGCN)。该模型可同时关注骨骼点与不同身体部位,构建更优拓扑结构,显著提升识别准确率。利用超图表示骨骼点间的类别关系,并融入图卷积网络以建模骨骼点间的高阶关系,增强特征表达。此外,提出的超图注意力模块与超图卷积模块分别在时序和通道维度优化拓扑建模,进一步提升特征表示能力。在三个常用数据集上进行了广泛实验,结果表明所提方法在与当前最先进的基于骨骼的方法对比中表现最优。
原文摘要 · Abstract (English)
In recent years, action recognition has received much attention and wide application due to its important role in video understanding. Most of the researches on action recognition methods focused on improving the performance via various deep learning methods rather than the classification of skeleton points. The topological modeling between skeleton points and body parts was seldom considered. Although some studies have used a data-driven approach to classify the topology of the skeleton point, the nature of the skeleton point in terms of kinematics has not been taken into consideration. Therefore, in this paper, we draw on the theory of kinematics to adapt the topological relations of the skeleton point and propose a topological relation classification based on body parts and distance from core of body. To synthesize these topological relations for action recognition, we propose a novel Hypergraph Fusion Graph Convolutional Network (HFGCN). In particular, the proposed model is able to focus on the human skeleton points and the different body parts simultaneously, and thus construct the topology, which improves the recognition accuracy obviously. We use a hypergraph to represent the categorical relationships of these skeleton points and incorporate the hypergraph into a graph convolution network to model the higher-order relationships among the skeleton points and enhance the feature representation of the network. In addition, our proposed hypergraph attention module and hypergraph graph convolution module optimize topology modeling in temporal and channel dimensions, respectively, to further enhance the feature representation of the network. We conducted extensive experiments on three widely used datasets.The results validate that our proposed method can achieve the best performance when compared with the state-of-the-art skeleton-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。