通过动态拓扑建模提升骨骼手势识别精度
DSTSA-GCN: Advancing Skeleton-Based Gesture Recognition with Semantic-Aware Spatio-Temporal Topology Modeling
- 分通道与帧级独立建模,捕捉动态骨骼结构变化
- 在多个基准数据集上达到最优性能,显著优于现有方法
- 适合关注骨骼动作识别与图卷积改进的研究者
图卷积网络(GCN)因其对骨架数据中时空依赖关系的建模能力,已成为基于骨架的动作与手势识别的强大工具。然而,现有基于GCN的方法存在两大局限:(1)缺乏有效建模动态骨骼运动中的时空拓扑结构;(2)难以捕捉超出局部关节连接的多尺度结构关系。为此,本文提出一种新框架——动态空间-时间语义感知图卷积网络(DSTSA-GCN)。该框架引入三个核心模块:分通道图卷积(GC-GC)、分帧图卷积(GT-GC)和多尺度时间卷积(MS-TCN)。GC-GC与GT-GC并行运行,分别建模通道特异性和帧特异性相关性,实现对时变拓扑的鲁棒学习;两者均采用分组策略以自适应捕获多尺度结构关系。同时,MS-TCN通过具有不同感受野的分组时间卷积增强时间建模能力。大量实验表明,DSTSA-GCN显著提升了GCN的拓扑建模能力,在多个基准数据集(包括SHREC17 Track、DHG-14/28、NTU-RGB+D 和 NTU-RGB+D-120)上实现了最先进的手势与动作识别性能。
原文摘要 · Abstract (English)
Graph convolutional networks (GCNs) have emerged as a powerful tool for skeleton-based action and gesture recognition, thanks to their ability to model spatial and temporal dependencies in skeleton data. However, existing GCN-based methods face critical limitations: (1) they lack effective spatio-temporal topology modeling that captures dynamic variations in skeletal motion, and (2) they struggle to model multiscale structural relationships beyond local joint connectivity. To address these issues, we propose a novel framework called Dynamic Spatial-Temporal Semantic Awareness Graph Convolutional Network (DSTSA-GCN). DSTSA-GCN introduces three key modules: Group Channel-wise Graph Convolution (GC-GC), Group Temporal-wise Graph Convolution (GT-GC), and Multi-Scale Temporal Convolution (MS-TCN). GC-GC and GT-GC operate in parallel to independently model channel-specific and frame-specific correlations, enabling robust topology learning that accounts for temporal variations. Additionally, both modules employ a grouping strategy to adaptively capture multiscale structural relationships. Complementing this, MS-TCN enhances temporal modeling through group-wise temporal convolutions with diverse receptive fields. Extensive experiments demonstrate that DSTSA-GCN significantly improves the topology modeling capabilities of GCNs, achieving state-of-the-art performance on benchmark datasets for gesture and action recognition, including SHREC17 Track, DHG-14\/28, NTU-RGB+D, and NTU-RGB+D-120.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。