arXiv:2509.23376cs.CV2025-09中稿 · PRCV 2025被引 1

用无标注的RGB-D数据,把2D姿态标注迁移到3D,省去人工标3D关键点。

UniPose: Unified Cross-modality Pose Prior Propagation towards RGB-D data for Weakly Supervised 3D Human Pose Estimation

  • 通过自监督学习,将2D数据标注迁移到3D点云,无需人工标注3D关键点。
  • 在CMU Panoptic和ITOP上表现接近全监督方法,且在大量无标签数据下更优。
  • 新提出的3D提升方法对2D误差更鲁棒,适合实际部署场景。

本文提出UniPose,一种统一的跨模态姿态先验传播方法,用于弱监督三维人体姿态估计(HPE),基于未标注的单视角RGB-D序列(含RGB、深度图和点云)。UniPose通过自监督学习,将大规模RGB数据集(如MS COCO)中的2D姿态标注迁移到3D空间,避免了繁琐的3D关键点标注。该方法不依赖多视角相机标定或合成到真实数据的转换问题。训练时,使用现成的2D姿态估计作为弱监督信号,结合身体对称性与关节运动等时空约束。2D到3D的反投影损失与跨模态交互进一步增强效果。以点云网络输出的3D HPE结果作为伪真值,采用锚点到关节的预测方法实现对RGB与深度网络的3D提升,相比现有方法更具鲁棒性。在CMU Panoptic和ITOP数据集上的实验表明,UniPose性能接近全监督方法。引入大规模无标签数据(如NTU RGB+D 60)后,在复杂条件下表现更佳,展现出实际应用潜力。所提3D提升方法也达到当前最优水平。

原文摘要 · Abstract (English)

In this paper, we present UniPose, a unified cross-modality pose prior propagation method for weakly supervised 3D human pose estimation (HPE) using unannotated single-view RGB-D sequences (RGB, depth, and point cloud data). UniPose transfers 2D HPE annotations from large-scale RGB datasets (e.g., MS COCO) to the 3D domain via self-supervised learning on easily acquired RGB-D sequences, eliminating the need for labor-intensive 3D keypoint annotations. This approach bridges the gap between 2D and 3D domains without suffering from issues related to multi-view camera calibration or synthetic-to-real data shifts. During training, UniPose leverages off-the-shelf 2D pose estimations as weak supervision for point cloud networks, incorporating spatial-temporal constraints like body symmetry and joint motion. The 2D-to-3D back-projection loss and cross-modality interaction further enhance this process. By treating the point cloud network's 3D HPE results as pseudo ground truth, our anchor-to-joint prediction method performs 3D lifting on RGB and depth networks, making it more robust against inaccuracies in 2D HPE results compared to state-of-the-art methods. Experiments on CMU Panoptic and ITOP datasets show that UniPose achieves comparable performance to fully supervised methods. Incorporating large-scale unlabeled data (e.g., NTU RGB+D 60) enhances its performance under challenging conditions, demonstrating its potential for practical applications. Our proposed 3D lifting method also achieves state-of-the-art results.

3D姿态估计弱监督跨模态点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。