arXiv:2501.05066cs.CVcs.AI2025-01被引 1

引入交互物体信息提升骨骼动作识别准确率

Improving Skeleton-based Action Recognition with Interactive Object Information

  • 用可变图结构融合人体骨骼与交互物体节点
  • 在NTU数据集上达到96.7%(跨用户)和99.2%(跨视角)准确率
  • 提出新数据集与抗过拟合增强方法,适合动作识别研究者

人体骨骼信息在基于骨骼的动作识别中至关重要,能高效描述人体姿态。然而现有方法主要关注骨骼信息,忽视人与物体的交互,导致涉及物体交互的动作识别性能不佳。本文提出一种新框架,通过引入物体节点补充缺失的交互物体信息,并设计时空可变图卷积网络(ST-VGCN)有效建模含物体节点的可变图(VG)。为验证交互物体信息的作用,采用简单自训练方法构建新数据集JXGC 24及扩展数据集NTU RGB+D+Object 60,包含超过200万额外物体节点。同时设计可变图构建方法以支持节点数量变化。首次探索引入额外物体信息带来的过拟合问题,提出基于可变图的数据增强方法——随机节点攻击(Random Node Attack)。在模型结构上,引入两种融合模块CAF与WNPool,以及新型节点平衡损失(Node Balance Loss),有效融合并平衡骨骼与物体节点信息。所提方法在多个骨架动作识别基准上超越现有最优结果,在NTU RGB+D 60跨用户划分上准确率达96.7%,跨视角划分上达99.2%。

原文摘要 · Abstract (English)

Human skeleton information is important in skeleton-based action recognition, which provides a simple and efficient way to describe human pose. However, existing skeleton-based methods focus more on the skeleton, ignoring the objects interacting with humans, resulting in poor performance in recognizing actions that involve object interactions. We propose a new action recognition framework introducing object nodes to supplement absent interactive object information. We also propose Spatial Temporal Variable Graph Convolutional Networks (ST-VGCN) to effectively model the Variable Graph (VG) containing object nodes. Specifically, in order to validate the role of interactive object information, by leveraging a simple self-training approach, we establish a new dataset, JXGC 24, and an extended dataset, NTU RGB+D+Object 60, including more than 2 million additional object nodes. At the same time, we designe the Variable Graph construction method to accommodate a variable number of nodes for graph structure. Additionally, we are the first to explore the overfitting issue introduced by incorporating additional object information, and we propose a VG-based data augmentation method to address this issue, called Random Node Attack. Finally, regarding the network structure, we introduce two fusion modules, CAF and WNPool, along with a novel Node Balance Loss, to enhance the comprehensive performance by effectively fusing and balancing skeleton and object node information. Our method surpasses the previous state-of-the-art on multiple skeleton-based action recognition benchmarks. The accuracy of our method on NTU RGB+D 60 cross-subject split is 96.7\%, and on cross-view split, it is 99.2\%.

动作识别图神经网络物体交互骨骼分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。