arXiv:2409.09444cs.CV2024-09被引 4

融合肢体微动与姿态整体结构,提升3D动作识别精度

KAN-HyperpointNet for Point Cloud Sequence-Based 3D Human Action Recognition

  • 用D-Hyperpoint同时捕捉局部运动与全局姿态
  • 结合KAN网络实现更优时空交互,准确率超越现有方法
  • 适合需要精细动作理解的场景,如医疗康复分析

基于点云序列的3D动作识别已取得显著性能与效率。然而,现有方法难以兼顾肢体微动的精确性与姿态宏观结构的完整性,导致动作推理中关键信息丢失。为此,我们提出D-Hyperpoint,一种通过D-Hyperpoint嵌入模块生成的新数据类型,可同时封装区域瞬时运动与全局静态姿态,有效概括每一时刻的单元人体动作。此外,我们设计了D-Hyperpoint KANsMixer模块,递归应用于分组后的D-Hyperpoints,以学习动作判别信息,并创新性地引入柯尔莫哥洛夫-阿诺德网络(KAN),增强D-Hyperpoints内部的时空交互。最后,提出KAN-HyperpointNet,一种时空解耦的3D动作识别网络架构。在MSR Action3D和NTU-RGB+D 60两个公开数据集上的大量实验表明,该方法达到当前最优性能。

原文摘要 · Abstract (English)

Point cloud sequence-based 3D action recognition has achieved impressive performance and efficiency. However, existing point cloud sequence modeling methods cannot adequately balance the precision of limb micro-movements with the integrity of posture macro-structure, leading to the loss of crucial information cues in action inference. To overcome this limitation, we introduce D-Hyperpoint, a novel data type generated through a D-Hyperpoint Embedding module. D-Hyperpoint encapsulates both regional-momentary motion and global-static posture, effectively summarizing the unit human action at each moment. In addition, we present a D-Hyperpoint KANsMixer module, which is recursively applied to nested groupings of D-Hyperpoints to learn the action discrimination information and creatively integrates Kolmogorov-Arnold Networks (KAN) to enhance spatio-temporal interaction within D-Hyperpoints. Finally, we propose KAN-HyperpointNet, a spatio-temporal decoupled network architecture for 3D action recognition. Extensive experiments on two public datasets: MSR Action3D and NTU-RGB+D 60, demonstrate the state-of-the-art performance of our method.

3D动作识别点云序列KAN网络姿态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。