用事件流的时空特性提升人体姿态估计效率与精度
Exploiting Spatiotemporal Properties for Efficient Event-Driven Human Pose Estimation
- 基于点云框架,利用事件流的时空特性建模
- 在DHP19上平均降低4%的MPJPE误差
- 适合追求低延迟高精度姿态估计的研究者
人体姿态估计旨在预测身体关键点以分析运动。目前多数方法依赖传统RGB相机,而事件相机具有高时间分辨率和低延迟,能在挑战性条件下实现鲁棒估计,为姿态估计带来新可能。然而,现有方法多将事件流转为密集事件帧,增加计算开销并损失事件信号的时间分辨率。本文提出一种基于点云框架的方法,充分利用事件流的时空特性,在保持计算高效的同时提升姿态估计性能。设计了事件时间切片卷积模块以捕捉短时依赖,并结合事件切片序列化模块实现结构化时序建模。此外,提出边缘增强型点云事件表示,强化稀疏事件条件下的空间边缘信息。在DHP19数据集上的实验表明,所提方法在三种代表性点云骨干网络(PointNet、DGCNN、Point Transformer)上均持续提升性能,平均MPJPE降低4%。
原文摘要 · Abstract (English)
Human pose estimation focuses on predicting body keypoints to analyze human motion. Currently, most pose estimation tasks rely on conventional RGB cameras. In contrast, event cameras provide high temporal resolution and low latency, enabling robust estimation under challenging conditions and opening up new possibilities for pose estimation. However, most existing methods convert event streams into dense event frames, which adds extra computation and sacrifices the high temporal resolution of the event signal. In this work, we aim to exploit the spatiotemporal properties of event streams based on point cloud-based framework, designed to enhance human pose estimation performance while maintaining computational efficiency. We design Event Temporal Slicing Convolution module to capture short-term dependencies across event slices, and combine it with Event Slice Sequencing module for structured temporal modeling. We further propose an edge-enhanced point cloud-based event representation to enhance spatial edge information under sparse event conditions to further improve performance. Experiments on the DHP19 dataset show our proposed method consistently improves performance across three representative point cloud backbones: PointNet, DGCNN, and Point Transformer, with an average MPJPE reduction of 4%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。