首个骨架动作理解基础模型,可通用解决多种动作识别任务。
Foundation Model for Skeleton-Based Human Action Understanding
- 构建统一的密集骨架表征学习框架,融合时空编码与多粒度特征去相关。
- 在25个基准上超越现有方法,尤其在密集预测任务中提升显著。
- 适合研究动作理解、机器人控制及跨任务泛化方向的学者参考。
人体动作理解是智能运动感知的基础。骨架作为与模态和设备无关的表示形式,在人形机器人控制与交互中有广泛应用前景。然而,现有方法普遍缺乏可扩展性与泛化能力,尚无适用于多样化动作理解任务的骨架基础模型。本文提出统一的骨架密集表征学习(USDRL)框架,包含基于Transformer的时空稠密编码器(DSTE)、多粒度特征去相关(MG-FD)与多视角一致性训练(MPCT)。DSTE通过双并行流学习时序动态与空间结构特征;MG-FD在时间、空间与实例维度协同进行特征去相关,减少冗余;MPCT结合多视图与多模态自监督一致性训练,增强高层语义学习并提升多模态特征表达。我们在9类骨架动作理解任务中的25个基准上进行实验,涵盖粗粒度预测、密集预测与迁移预测。结果表明,该方法显著优于当前最优模型,有望推动骨架动作理解研究的发展,并促进对密集预测任务的关注。
原文摘要 · Abstract (English)
Human action understanding serves as a foundational pillar in the field of intelligent motion perception. Skeletons serve as a modality- and device-agnostic representation for human modeling, and skeleton-based action understanding has potential applications in humanoid robot control and interaction. \RED{However, existing works often lack the scalability and generalization required to handle diverse action understanding tasks. There is no skeleton foundation model that can be adapted to a wide range of action understanding tasks}. This paper presents a Unified Skeleton-based Dense Representation Learning (USDRL) framework, which serves as a foundational model for skeleton-based human action understanding. USDRL consists of a Transformer-based Dense Spatio-Temporal Encoder (DSTE), Multi-Grained Feature Decorrelation (MG-FD), and Multi-Perspective Consistency Training (MPCT). The DSTE module adopts two parallel streams to learn temporal dynamic and spatial structure features. The MG-FD module collaboratively performs feature decorrelation across temporal, spatial, and instance domains to reduce dimensional redundancy and enhance information extraction. The MPCT module employs both multi-view and multi-modal self-supervised consistency training. The former enhances the learning of high-level semantics and mitigates the impact of low-level discrepancies, while the latter effectively facilitates the learning of informative multimodal features. We perform extensive experiments on 25 benchmarks across across 9 skeleton-based action understanding tasks, covering coarse prediction, dense prediction, and transferred prediction. Our approach significantly outperforms the current state-of-the-art methods. We hope that this work would broaden the scope of research in skeleton-based action understanding and encourage more attention to dense prediction tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。