arXiv:2603.02477cs.CV2026-03

基于骨骼数据的端到端几何神经网络,提升动作识别准确率。

E2E-GNet: An End-to-End Skeleton-based Geometric Deep Neural Network for Human Motion Recognition

  • 在非欧空间中通过几何变换层优化骨骼序列
  • 投影后仍保持几何特征,识别准确率显著提高
  • 适合需要高精度动作识别的应用场景

几何深度学习近年来在计算机视觉领域受到广泛关注,因其能捕捉非欧几里得空间中的数据有意义表示。为此,本文提出E2E-GNet,一种面向骨骼数据的人体动作识别端到端几何深度神经网络。为增强非欧空间中不同动作间的区分能力,E2E-GNet引入几何变换层,在该空间联合优化骨骼运动序列,并采用可微对数映射激活函数将其投影至线性空间。在此基础上,进一步设计了畸变感知优化层,限制投影引起的骨架形状失真,使网络得以保留判别性几何线索,实现更高动作识别率。通过消融实验及在五个跨三个领域的数据集上的广泛测试,证明E2E-GNet在更低计算成本下优于现有方法。

原文摘要 · Abstract (English)

Geometric deep learning has recently gained significant attention in the computer vision community for its ability to capture meaningful representations of data lying in a non-Euclidean space. To this end, we propose E2E-GNet, an end-to-end geometric deep neural network for skeleton-based human motion recognition. To enhance the discriminative power between different motions in the non-Euclidean space, E2E-GNet introduces a geometric transformation layer that jointly optimizes skeleton motion sequences on this space and applies a differentiable logarithm map activation to project them onto a linear space. Building on this, we further design a distortion-aware optimization layer that limits skeleton shape distortions caused by this projection, enabling the network to retain discriminative geometric cues and achieve a higher motion recognition rate. We demonstrate the impact of each layer through ablation studies and extensive experiments across five datasets spanning three domains show that E2E-GNet outperforms other methods with lower cost.

动作识别几何深度学习骨骼数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。