arXiv:2411.12560cs.CVcs.AI2024-11

通过人体对称性增强图卷积,提升骨骼动作识别精度

Topological Symmetry Enhanced Graph Convolution for Skeleton-Based Action Recognition

  • 引入人体对称性先验,分通道学习动态拓扑结构
  • 在NTU RGB+D 120上跨主体/跨设置准确率达90.0%/91.1%
  • 参数量仅110万,适合低资源部署场景

基于骨架的动作识别得益于图卷积网络(GCNs)的发展取得了显著进展。然而,现有方法多构建复杂的拓扑学习机制,忽视了人体固有的对称性。此外,使用固定感受野的时序卷积限制了其捕捉时间依赖关系的能力。为此,本文提出一种新型拓扑对称性增强图卷积(TSE-GC),在不同通道划分下实现差异化拓扑学习,并融入拓扑对称性感知;同时设计多分支可变形时序卷积(MBDTC),引入可变形建模思想,获得更灵活的感受野和更强的时间依赖建模能力。将TSE-GC与MBDTC结合,所提模型TSE-GCN在三个大规模数据集(NTU RGB+D、NTU RGB+D 120、NW-UCLA)上达到媲美前沿方法的性能,且参数更少。在NTU RGB+D 120的跨主体和跨设置评估中,准确率分别达到90.0%和91.1%,单流模型参数量为1.1M,计算量为1.38 GFLOPS。

原文摘要 · Abstract (English)

Skeleton-based action recognition has achieved remarkable performance with the development of graph convolutional networks (GCNs). However, most of these methods tend to construct complex topology learning mechanisms while neglecting the inherent symmetry of the human body. Additionally, the use of temporal convolutions with certain fixed receptive fields limits their capacity to effectively capture dependencies in time sequences. To address the issues, we (1) propose a novel Topological Symmetry Enhanced Graph Convolution (TSE-GC) to enable distinct topology learning across different channel partitions while incorporating topological symmetry awareness and (2) construct a Multi-Branch Deformable Temporal Convolution (MBDTC) for skeleton-based action recognition. The proposed TSE-GC emphasizes the inherent symmetry of the human body while enabling efficient learning of dynamic topologies. Meanwhile, the design of MBDTC introduces the concept of deformable modeling, leading to more flexible receptive fields and stronger modeling capacity of temporal dependencies. Combining TSE-GC with MBDTC, our final model, TSE-GCN, achieves competitive performance with fewer parameters compared with state-of-the-art methods on three large datasets, NTU RGB+D, NTU RGB+D 120, and NW-UCLA. On the cross-subject and cross-set evaluations of NTU RGB+D 120, the accuracies of our model reach 90.0\% and 91.1\%, with 1.1M parameters and 1.38 GFLOPS for one stream.

动作识别图卷积对称性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。