用点云生成3D人体姿态,解决遮挡与数据不足难题
Point2Pose: A Generative Framework for 3D Human Pose Estimation with Multi-View Point Cloud Dataset
- 通过时空点云编码器+注意力生成回归器建模姿态分布
- 在MVPose3D等数据集上显著优于基线模型
- 适合做多视角点云姿态估计的科研与工业应用
我们提出一种新型生成式3D人体姿态估计方法。由于人体结构复杂、关节自遮挡及真实世界运动数据规模不足,3D人体姿态估计面临诸多挑战。为此,我们引入Point2Pose框架,通过序列点云和姿态历史条件建模人体姿态分布。具体地,采用时空点云编码器与姿态特征编码器提取关节点特征,再通过基于注意力的生成回归器输出姿态。此外,我们构建了大规模室内多模态数据集MVPose3D,包含非平凡人体动作的IMU数据、密集多视角点云及RGB图像。实验表明,所提方法在多个数据集上均优于基线模型,性能更优。
原文摘要 · Abstract (English)
We propose a novel generative approach for 3D human pose estimation. 3D human pose estimation poses several key challenges due to the complex geometry of the human body, self-occluding joints, and the requirement for large-scale real-world motion datasets. To address these challenges, we introduce Point2Pose, a framework that effectively models the distribution of human poses conditioned on sequential point cloud and pose history. Specifically, we employ a spatio-temporal point cloud encoder and a pose feature encoder to extract joint-wise features, followed by an attention-based generative regressor. Additionally, we present a large-scale indoor dataset MVPose3D, which contains multiple modalities, including IMU data of non-trivial human motions, dense multi-view point clouds, and RGB images. Experimental results show that the proposed method outperforms the baseline models, demonstrating its superior performance across various datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。