解决不同骨骼结构动作识别难题,让机器人跨平台理解动作。
Heterogeneous Skeleton-Based Action Representation Learning

- 用提示词统一异构骨骼数据,先转三维再生成标准骨架
- 在三个数据集上准确率超现有方法,最高达93.2%
- 适合开发适配多型号机器人的通用动作识别系统
基于骨骼的人体动作识别近年来受到广泛关注,因其应用场景多样。由于骨骼数据来源不同,天然存在异构性。以往工作忽视了骨骼的异构性,仅针对同构骨骼构建模型。本文提出异构骨骼动作表征学习框架,重点处理关节数量和拓扑结构不同的骨骼数据。该框架包含两个核心模块:异构骨骼处理与统一表征学习。前者通过辅助网络将二维骨骼转化为三维,并利用骨骼特异性提示构建统一骨架;还设计了语义运动编码模态,以挖掘骨骼中的语义信息。后者模块采用共享主干网络,对不同异构骨骼进行统一动作表征学习。在NTU-60、NTU-120和PKU-MMD II数据集上的大量实验表明,该方法在多种动作理解任务中均有效。本方法可应用于具有不同人形结构的机器人动作识别。
原文摘要 · Abstract (English)
Skeleton-based human action recognition has received widespread attention in recent years due to its diverse range of application scenarios. Due to the different sources of human skeletons, skeleton data naturally exhibit heterogeneity. The previous works, however, overlook the heterogeneity of human skeletons and solely construct models tailored for homogeneous skeletons. This work addresses the challenge of heterogeneous skeleton-based action representation learning, specifically focusing on processing skeleton data that varies in joint dimensions and topological structures. The proposed framework comprises two primary components: heterogeneous skeleton processing and unified representation learning. The former first converts two-dimensional skeleton data into three-dimensional skeleton via an auxiliary network, and then constructs a prompted unified skeleton using skeleton-specific prompts. We also design an additional modality named semantic motion encoding to harness the semantic information within skeletons. The latter module learns a unified action representation using a shared backbone network that processes different heterogeneous skeletons. Extensive experiments on the NTU-60, NTU-120, and PKU-MMD II datasets demonstrate the effectiveness of our method in various tasks of action understanding. Our approach can be applied to action recognition in robots with different humanoid structures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。