arXiv:2501.02593cs.CVcs.AI2025-01中稿 · the Companion Proc…被引 7

对比传统骨架与运动增强骨架在动作识别中的表现

Evolving Skeletons: Motion Dynamics in Action Recognition

  • 用图和超图模型比较静态骨架与运动注入骨架
  • 在NTU-60和NTU-120上验证运动增强提升动态表征能力
  • 揭示运动增强骨架的潜力与当前应用瓶颈

基于骨架的动作识别因其轻量级时空信息表达而受到广泛关注。现有方法多采用图模型处理骨架序列,其中每个姿态以围绕人体物理连接结构的骨骼图为单位。在此类方法中,时空图卷积网络(ST-GCN)已成为广泛应用的框架。相比之下,超图模型如Hyperformer能捕捉更高阶相关性,提供更丰富的关节交互表示。近期提出的Taylor Videos通过嵌入运动概念生成运动增强骨架序列,为理解骨架动作提供了新视角。本文在NTU-60和NTU-120数据集上,使用ST-GCN与Hyperformer对传统骨架序列与Taylor变换后的骨架进行综合评估,比较了骨骼图与超图表示,分析静态姿态与运动注入姿态的表现差异。研究结果揭示了Taylor变换骨架在提升运动动态建模方面的优势,同时指出了其尚未充分挖掘的挑战。该工作强调了构建创新骨架建模技术以有效处理富含运动信息数据的必要性,推动动作识别领域发展。

原文摘要 · Abstract (English)

Skeleton-based action recognition has gained significant attention for its ability to efficiently represent spatiotemporal information in a lightweight format. Most existing approaches use graph-based models to process skeleton sequences, where each pose is represented as a skeletal graph structured around human physical connectivity. Among these, the Spatiotemporal Graph Convolutional Network (ST-GCN) has become a widely used framework. Alternatively, hypergraph-based models, such as the Hyperformer, capture higher-order correlations, offering a more expressive representation of complex joint interactions. A recent advancement, termed Taylor Videos, introduces motion-enhanced skeleton sequences by embedding motion concepts, providing a fresh perspective on interpreting human actions in skeleton-based action recognition. In this paper, we conduct a comprehensive evaluation of both traditional skeleton sequences and Taylor-transformed skeletons using ST-GCN and Hyperformer models on the NTU-60 and NTU-120 datasets. We compare skeletal graph and hypergraph representations, analyzing static poses against motion-injected poses. Our findings highlight the strengths and limitations of Taylor-transformed skeletons, demonstrating their potential to enhance motion dynamics while exposing current challenges in fully using their benefits. This study underscores the need for innovative skeletal modelling techniques to effectively handle motion-rich data and advance the field of action recognition.

动作识别骨架建模运动动态图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。