arXiv:2411.08128cs.CV2024-11被引 68

通过改进相机参数与身体关键点,提升单目图像人体三维重建精度。

CameraHMR: Aligning People with Perspective

  • 用新模型预测视角,更准地估计相机参数。
  • 结合密集表面关键点检测,生成更真实的人体形状。
  • 适合做3D人体重建的科研人员参考使用。

针对单目图像中3D人体姿态与形状估计的挑战,本文提出两个关键改进。首先,构建一个基于含人图像数据集训练的视野预测模型(HumanFoV),用于估计相机内参,并在4D-Humans数据集中引入完整透视相机模型,提升伪真值(pGT)质量。其次,2D关节对3D体型约束不足导致结果平均化,为此利用BEDLAM数据集训练密集表面关键点检测器,应用于4D-Humans并改进SMPLify拟合过程,显著提升体型真实性。最后,将估计的相机参数融入HMR2.0架构,迭代优化模型训练与拟合,获得更准确的伪真值和新模型CameraHMR,达到当前最优性能。代码与伪真值已开放供研究使用。

原文摘要 · Abstract (English)

We address the challenge of accurate 3D human pose and shape estimation from monocular images. The key to accuracy and robustness lies in high-quality training data. Existing training datasets containing real images with pseudo ground truth (pGT) use SMPLify to fit SMPL to sparse 2D joint locations, assuming a simplified camera with default intrinsics. We make two contributions that improve pGT accuracy. First, to estimate camera intrinsics, we develop a field-of-view prediction model (HumanFoV) trained on a dataset of images containing people. We use the estimated intrinsics to enhance the 4D-Humans dataset by incorporating a full perspective camera model during SMPLify fitting. Second, 2D joints provide limited constraints on 3D body shape, resulting in average-looking bodies. To address this, we use the BEDLAM dataset to train a dense surface keypoint detector. We apply this detector to the 4D-Humans dataset and modify SMPLify to fit the detected keypoints, resulting in significantly more realistic body shapes. Finally, we upgrade the HMR2.0 architecture to include the estimated camera parameters. We iterate model training and SMPLify fitting initialized with the previously trained model. This leads to more accurate pGT and a new model, CameraHMR, with state-of-the-art accuracy. Code and pGT are available for research purposes.

3D人体重建单目估计相机建模数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。