arXiv:2410.23806cs.CV2024-10被引 1

用时空相对注意力网络提升骨骼动作识别长程依赖建模能力

Human Action Recognition (HAR) Using Skeleton-based Spatial Temporal Relative Transformer Network: ST-RTR

  • 设计关节与中继节点结构,打破骨架拓扑限制以捕捉长程关联
  • 在NTU RGB+D 60/120和UAV-Human数据集上分别提升准确率2.11%/2.54%
  • 适合关注骨骼动作识别中长时序建模的开发者与研究者

人体动作识别(HAR)是人机交互中的重要研究方向,用于监测老年人及残障人士的健康状况。近年来,基于骨骼数据的HAR受到广泛关注,因其能有效应对姿态变化、体型差异、视角变换及复杂背景等问题。虽然时空图卷积网络(ST-GCN)可自动学习骨骼序列中的时空模式,但其感受野有限,仅适用于短程关联,难以捕捉人类动作的长程依赖关系。为此,本文提出空间-时间相对变压器网络(ST-RTR),引入关节与中继节点,实现网络内高效通信与数据传输,打破原有骨架的空间与时间拓扑结构,从而更有效地理解长程动作。此外,将ST-RTR与融合模型结合进一步提升性能。在三个基于骨骼的HAR基准数据集(NTU RGB+D 60、NTU RGB+D 120、UAV-Human)上的实验表明:在NTU RGB+D 60上,分类精度(CS)和视图验证(CV)分别提升2.11%和1.45%;在NTU RGB+D 120上分别提升1.25%和1.05%;在UAV-Human数据集上准确率提升2.54%。结果表明,所提出的ST-RTR模型显著优于标准ST-GCN方法。

原文摘要 · Abstract (English)

Human Action Recognition (HAR) is an interesting research area in human-computer interaction used to monitor the activities of elderly and disabled individuals affected by physical and mental health. In the recent era, skeleton-based HAR has received much attention because skeleton data has shown that it can handle changes in striking, body size, camera views, and complex backgrounds. One key characteristic of ST-GCN is automatically learning spatial and temporal patterns from skeleton sequences. It has some limitations, as this method only works for short-range correlation due to its limited receptive field. Consequently, understanding human action requires long-range interconnection. To address this issue, we developed a spatial-temporal relative transformer ST-RTR model. The ST-RTR includes joint and relay nodes, which allow efficient communication and data transmission within the network. These nodes help to break the inherent spatial and temporal skeleton topologies, which enables the model to understand long-range human action better. Furthermore, we combine ST-RTR with a fusion model for further performance improvements. To assess the performance of the ST-RTR method, we conducted experiments on three skeleton-based HAR benchmarks: NTU RGB+D 60, NTU RGB+D 120, and UAV-Human. It boosted CS and CV by 2.11 % and 1.45% on NTU RGB+D 60, 1.25% and 1.05% on NTU RGB+D 120. On UAV-Human datasets, accuracy improved by 2.54%. The experimental outcomes explain that the proposed ST-RTR model significantly improves action recognition associated with the standard ST-GCN method.

动作识别骨架分析注意力机制时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。