arXiv:2412.13454cs.CVcs.AI2024-12AAAI被引 5

通过建模点云密度特性,提升低质量激光雷达下人体姿态估计精度。

Pre-training a Density-Aware Pose Transformer for Robust LiDAR-based 3D Human Pose Estimation

  • 设计密度感知的变换器,利用关节锚点与交换模块提取多密度点云信息。
  • 在Waymo数据集上比LPFormer降低10.0mm平均MPJPE,SLOPER4D上降低20.7mm。
  • 适合关注激光雷达人体姿态估计、点云增强与自监督预训练的研究者。

随着自动驾驶快速发展,基于激光雷达的3D人体姿态估计(3D HPE)日益成为研究热点。然而,由于激光雷达点云存在噪声和稀疏性,鲁棒的人体姿态估计仍具挑战。现有方法多依赖时序信息、多模态融合或SMPL优化来修正偏差。本文仅通过建模低质量点云的内在特性,提出一种简单而有效的密度感知姿态变换器(DAPT),以获得稳定的关键点表征。通过一组关节锚点与精心设计的交换模块,从不同密度点云中提取有效信息,并使用1D热图精确表示关键点位置。其次,提出全面的激光雷达人体合成与增强方法用于模型预训练,使模型具备更好的人体先验知识。通过随机采样人体位置与朝向,并引入激光级掩码模拟遮挡,显著提升点云多样性。在多个数据集上的大量实验表明,本方法在所有场景下均达到当前最优性能。尤其在Waymo数据集上,相比LPFormer,平均MPJPE降低10.0mm;在SLOPER4D上,相比PRN,平均MPJPE降低20.7mm。

原文摘要 · Abstract (English)

With the rapid development of autonomous driving, LiDAR-based 3D Human Pose Estimation (3D HPE) is becoming a research focus. However, due to the noise and sparsity of LiDAR-captured point clouds, robust human pose estimation remains challenging. Most of the existing methods use temporal information, multi-modal fusion, or SMPL optimization to correct biased results. In this work, we try to obtain sufficient information for 3D HPE only by modeling the intrinsic properties of low-quality point clouds. Hence, a simple yet powerful method is proposed, which provides insights both on modeling and augmentation of point clouds. Specifically, we first propose a concise and effective density-aware pose transformer (DAPT) to get stable keypoint representations. By using a set of joint anchors and a carefully designed exchange module, valid information is extracted from point clouds with different densities. Then 1D heatmaps are utilized to represent the precise locations of the keypoints. Secondly, a comprehensive LiDAR human synthesis and augmentation method is proposed to pre-train the model, enabling it to acquire a better human body prior. We increase the diversity of point clouds by randomly sampling human positions and orientations and by simulating occlusions through the addition of laser-level masks. Extensive experiments have been conducted on multiple datasets, including IMU-annotated LidarHuman26M, SLOPER4D, and manually annotated Waymo Open Dataset v2.0 (Waymo), HumanM3. Our method demonstrates SOTA performance in all scenarios. In particular, compared with LPFormer on Waymo, we reduce the average MPJPE by $10.0mm$. Compared with PRN on SLOPER4D, we notably reduce the average MPJPE by $20.7mm$.

3D姿态估计激光雷达点云处理变换器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。