用自监督方法提取人体动作的层次化语义单元,提升行为建模效果。
Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movements

- 将人体动作分解为原子级关节运动(Action Atoms)和时间组合的语义模式(Action Motifs)
- 通过掩码预测任务在潜空间中学习,自动发现可复用的动作片段
- 适用于动作识别、预测与插值,尤其适合有遮挡的多视角视频分析
有效的人类行为建模需要利用动作的组合性。本文提出一种层次化表示,包含捕捉基本关节运动的Action Atoms,以及由其时序组合形成的、编码跨不同动作的相似身体运动的Action Motifs。我们提出A4Mer,一种嵌套的潜在Transformer,可从3D姿态数据中完全自监督地学习该表示。A4Mer将3D姿态序列切分为可变长度段,并将每段表示为单一潜变量(Action Atoms)。通过自底向上的表征学习,这些原子组合出有意义的时间模式(Action Motifs),自然形成可复用的语义动作段。A4Mer通过统一的掩码令牌预测预训练任务实现此目标。我们还构建了大规模多视角人体行为数据集Action Motif Dataset (AMD),包含完整SMPL标注。创新性地将摄像头安装于脚部,克服频繁且严重的身体遮挡,实现帧级标注。实验表明,A4Mer能有效提取有意义的Action Motifs,显著提升动作识别、运动预测与运动插值等任务性能。
原文摘要 · Abstract (English)
Effective human behavior modeling requires a representation of the human body movement that capitalizes on its compositionality. We propose a hierarchical representation consisting of Action Atoms that capture the atomic joint movements and Action Motifs that are formed by their temporal compositions and encode similar body movements found across different overall human actions. We derive A4Mer, a nested latent Transformer to learn this hierarchical representation from human pose data in a fully self-supervised manner. A4Mer splits a 3D pose sequence into variable-length segments and represents each segment as a single latent token (Action Atoms). Through bottom-up representation learning, temporal patterns composed of these Action Atoms, which capture meaningful temporal spans of reusable, semantic segments of body movements, naturally emerge (Action Motifs). A4Mer achieves this with a unified pretext task of masked token prediction in their respective latent spaces. We also introduce Action Motif Dataset (AMD), a large-scale dataset of multi-view human behavior videos with full SMPL annotations. We introduce a novel use of cameras by mounting them on the feet to achieve their frame-wise annotations despite frequent and heavy body occlusions. Experimental results demonstrate the effectiveness of A4Mer for extracting meaningful Action Motifs, which significantly benefit human behavior modeling tasks including action recognition, motion prediction, and motion interpolation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。