用语义引导遮蔽,学习3D骨骼的高阶运动特征,提升动作识别效果。
MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
- 基于相对运动的Grad-CAM指导关节遮蔽,聚焦语义丰富的时序区域。
- 以速度与加速度联合为重建目标,捕捉多阶运动模式,准确率提升显著。
- 适用于人机协作场景,尤其适合对复杂动作理解有要求的应用。
人体动作识别是智能机器人领域的重要任务,尤其在人机协作研究中具有关键意义。在自监督骨架动作识别中,基于掩码的重建范式通过遮蔽关节并从无标签数据中重建目标来学习骨架的空间结构与运动模式。然而,现有方法仅关注有限关节和低阶运动模式,限制了模型对复杂运动的理解能力。为此,本文提出MaskSem,一种用于学习3D混合高阶运动表示的语义引导遮蔽方法。该框架利用基于相对运动的Grad-CAM生成语义重要性图,指导关节遮蔽,从而定位最具语义信息的时序区域。同时,引入混合高阶运动作为重建目标,将低阶运动速度与高阶运动加速度共同用于重建,更全面地描述动态运动过程。在NTU60、NTU120和PKU-MMD数据集上的实验表明,结合基础Transformer,MaskSem显著提升了骨架动作识别性能,更适合应用于人机交互场景。
原文摘要 · Abstract (English)
Human action recognition is a crucial task for intelligent robotics, particularly within the context of human-robot collaboration research. In self-supervised skeleton-based action recognition, the mask-based reconstruction paradigm learns the spatial structure and motion patterns of the skeleton by masking joints and reconstructing the target from unlabeled data. However, existing methods focus on a limited set of joints and low-order motion patterns, limiting the model's ability to understand complex motion patterns. To address this issue, we introduce MaskSem, a novel semantic-guided masking method for learning 3D hybrid high-order motion representations. This novel framework leverages Grad-CAM based on relative motion to guide the masking of joints, which can be represented as the most semantically rich temporal orgions. The semantic-guided masking process can encourage the model to explore more discriminative features. Furthermore, we propose using hybrid high-order motion as the reconstruction target, enabling the model to learn multi-order motion patterns. Specifically, low-order motion velocity and high-order motion acceleration are used together as the reconstruction target. This approach offers a more comprehensive description of the dynamic motion process, enhancing the model's understanding of motion patterns. Experiments on the NTU60, NTU120, and PKU-MMD datasets show that MaskSem, combined with a vanilla transformer, improves skeleton-based action recognition, making it more suitable for applications in human-robot interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。