系统梳理2014-2024年多模态人体行为识别进展,涵盖主流方法与数据集。
A Comprehensive Methodological Survey of Human Activity Recognition Across Divers Data Modalities
- 按输入模态分类,综述机器学习与深度学习方法
- 对比分析单模态与多模态融合、协同学习框架性能
- 适合研究者快速掌握领域现状与未来方向
人体行为识别(HAR)系统旨在理解人类行为并为每个动作打标签,在计算机视觉中因应用广泛而备受关注。HAR可利用多种数据模态,如RGB图像与视频、骨骼数据、深度图、红外、点云、事件流、音频、加速度和雷达信号。每种模态提供独特且互补的信息,适用于不同场景。因此,大量研究探索了基于这些模态的各类方法。本文系统回顾2014至2024年HAR领域的最新进展,聚焦于按输入模态划分的机器学习(ML)与深度学习(DL)方法。涵盖单模态与多模态技术,重点分析融合型与协同学习框架。此外,还包括手工特征设计、人机交互识别及活动检测的进展。针对每种模态,详细描述常用数据集,并总结最新HAR系统在基准数据集上的比较结果。最后,提出深刻见解并建议有效的未来研究方向。
原文摘要 · Abstract (English)
Human Activity Recognition (HAR) systems aim to understand human behaviour and assign a label to each action, attracting significant attention in computer vision due to their wide range of applications. HAR can leverage various data modalities, such as RGB images and video, skeleton, depth, infrared, point cloud, event stream, audio, acceleration, and radar signals. Each modality provides unique and complementary information suited to different application scenarios. Consequently, numerous studies have investigated diverse approaches for HAR using these modalities. This paper presents a comprehensive survey of the latest advancements in HAR from 2014 to 2024, focusing on machine learning (ML) and deep learning (DL) approaches categorized by input data modalities. We review both single-modality and multi-modality techniques, highlighting fusion-based and co-learning frameworks. Additionally, we cover advancements in hand-crafted action features, methods for recognizing human-object interactions, and activity detection. Our survey includes a detailed dataset description for each modality and a summary of the latest HAR systems, offering comparative results on benchmark datasets. Finally, we provide insightful observations and propose effective future research directions in HAR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。