系统梳理3D骨骼动作识别的全流程,补全模型之外的关键环节。
3D Skeleton-Based Action Recognition: A Review
- 按任务拆解骨架动作识别,涵盖预处理、特征提取与时空建模
- 总结主流数据集与前沿模型,包括混合架构和大语言模型应用
- 适合想全面掌握该领域研究脉络的学者和工程师
基于骨骼表示的3D动作识别在计算机视觉中日益重要。以往综述多聚焦于模型设计,忽略了骨架动作识别中的基础步骤,导致对任务本质理解不足。为此,本文提出一个任务导向的综合框架,将任务分解为多个子任务,重点分析预处理环节如模态生成与数据增强,深入探讨特征提取与时空建模技术。除传统识别网络外,还涵盖了混合架构、Mamba模型、大语言模型(LLMs)及生成模型等最新进展。最后,全面梳理公开的3D骨骼数据集,并分析其上表现最优的算法。通过整合任务视角、子任务剖析与前沿技术,本综述为理解和推进该领域提供了清晰、系统的路线图。
原文摘要 · Abstract (English)
With the inherent advantages of skeleton representation, 3D skeleton-based action recognition has become a prominent topic in the field of computer vision. However, previous reviews have predominantly adopted a model-oriented perspective, often neglecting the fundamental steps involved in skeleton-based action recognition. This oversight tends to ignore key components of skeleton-based action recognition beyond model design and has hindered deeper, more intrinsic understanding of the task. To bridge this gap, our review aims to address these limitations by presenting a comprehensive, task-oriented framework for understanding skeleton-based action recognition. We begin by decomposing the task into a series of sub-tasks, placing particular emphasis on preprocessing steps such as modality derivation and data augmentation. The subsequent discussion delves into critical sub-tasks, including feature extraction and spatio-temporal modeling techniques. Beyond foundational action recognition networks, recently advanced frameworks such as hybrid architectures, Mamba models, large language models (LLMs), and generative models have also been highlighted. Finally, a comprehensive overview of public 3D skeleton datasets is presented, accompanied by an analysis of state-of-the-art algorithms evaluated on these benchmarks. By integrating task-oriented discussions, comprehensive examinations of sub-tasks, and an emphasis on the latest advancements, our review provides a fundamental and accessible structured roadmap for understanding and advancing the field of 3D skeleton-based action recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。