arXiv:2501.13066cs.CV2025-01综述被引 24

系统梳理动作识别技术,提出新分类框架,助力理解主流方法与未来方向。

SMART-Vision: Survey of Modern Action Recognition Techniques in Vision

  • 提出SMART-Vision新分类体系,融合多模态与混合架构思路。
  • 梳理从基础到前沿的算法演进路径,覆盖主流数据集与性能对比。
  • 聚焦开放环境下的动作识别挑战,适合研究者把握领域趋势。

人体动作识别(HAR)是计算机视觉中的关键挑战,需通过分析视频中个体运动的时空动态来识别复杂模式。这些模式存在于序列数据(如视频帧)中,对区分单帧下易混淆的动作至关重要。由于在机器人、监控、体育分析、医疗健康及自动驾驶等领域的广泛应用,HAR受到广泛关注。尽管已有多种分类体系,但常忽略混合方法,且未充分展示模型如何整合不同架构与模态。本文提出全新的SMART-Vision分类框架,揭示深度学习创新在HAR中的互补关系,推动超越传统分类的混合方法发展。本综述清晰勾勒出从基础工作到当前最先进系统的演进路线,指出新兴研究方向并讨论架构层面未解难题。同时详细列出各类方法所用的数据集,并探讨快速发展的开放域动作识别(Open-HAR)系统,该类系统在测试时面临未知新类别的挑战。

原文摘要 · Abstract (English)

Human Action Recognition (HAR) is a challenging domain in computer vision, involving recognizing complex patterns by analyzing the spatiotemporal dynamics of individuals' movements in videos. These patterns arise in sequential data, such as video frames, which are often essential to accurately distinguish actions that would be ambiguous in a single image. HAR has garnered considerable interest due to its broad applicability, ranging from robotics and surveillance systems to sports motion analysis, healthcare, and the burgeoning field of autonomous vehicles. While several taxonomies have been proposed to categorize HAR approaches in surveys, they often overlook hybrid methodologies and fail to demonstrate how different models incorporate various architectures and modalities. In this comprehensive survey, we present the novel SMART-Vision taxonomy, which illustrates how innovations in deep learning for HAR complement one another, leading to hybrid approaches beyond traditional categories. Our survey provides a clear roadmap from foundational HAR works to current state-of-the-art systems, highlighting emerging research directions and addressing unresolved challenges in discussion sections for architectures within the HAR domain. We provide details of the research datasets that various approaches used to measure and compare goodness HAR approaches. We also explore the rapidly emerging field of Open-HAR systems, which challenges HAR systems by presenting samples from unknown, novel classes during test time.

动作识别综述开放域视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。