arXiv:2510.05506cs.CV2025-10

用3D点云识别动作,支持深度传感器和单目估计数据

Human Action Recognition from Point Clouds over Time

  • 分步处理:分割人体、跟踪个体、划分身体部位
  • 融合法向量等特征,在NTU RGB-D 120上达89.3%准确率
  • 适合做点云动作识别的开发者,尤其关注真实场景应用

近期的人体动作识别研究主要集中在骨骼识别和基于视频的方法。随着消费级深度传感器和激光雷达设备的普及,利用密集3D数据进行动作识别成为可能,形成第三种路径。本文提出一种新方法,通过分割人体点云、追踪个体随时间变化,并进行身体部位分割,实现从3D视频中识别动作。该方法支持来自深度传感器及单目深度估计的点云输入。核心是结合点式技术与稀疏卷积网络的新型3D动作识别骨干网络,应用于体素化点云序列。实验引入表面法向量、颜色、红外强度及身体部位解析标签等辅助特征,提升识别精度。在NTU RGB-D 120数据集上的评估表明,该方法性能可媲美现有骨骼识别算法。采用传感器与估计深度数据的集成方案,在不同人物训练与测试条件下,达到89.3%的准确率,优于以往点云动作识别方法。

原文摘要 · Abstract (English)

Recent research into human action recognition (HAR) has focused predominantly on skeletal action recognition and video-based methods. With the increasing availability of consumer-grade depth sensors and Lidar instruments, there is a growing opportunity to leverage dense 3D data for action recognition, to develop a third way. This paper presents a novel approach for recognizing actions from 3D videos by introducing a pipeline that segments human point clouds from the background of a scene, tracks individuals over time, and performs body part segmentation. The method supports point clouds from both depth sensors and monocular depth estimation. At the core of the proposed HAR framework is a novel backbone for 3D action recognition, which combines point-based techniques with sparse convolutional networks applied to voxel-mapped point cloud sequences. Experiments incorporate auxiliary point features including surface normals, color, infrared intensity, and body part parsing labels, to enhance recognition accuracy. Evaluation on the NTU RGB- D 120 dataset demonstrates that the method is competitive with existing skeletal action recognition algorithms. Moreover, combining both sensor-based and estimated depth inputs in an ensemble setup, this approach achieves 89.3% accuracy when different human subjects are considered for training and testing, outperforming previous point cloud action recognition methods.

动作识别点云处理3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。