利用2D姿态检测器的不确定性提升3D人体网格生成的多样性与准确性
Utilizing Uncertainty in 2D Pose Detectors for Probabilistic 3D Human Mesh Recovery
- 通过热图分布监督3D预测,增强对模糊部位的合理性建模
- 在3DPW和EMDB上优于现有概率方法,尤其改善了不可见关节的估计
- 引入分割掩码减少无效样本,适合需要鲁棒3D重建的研究者
单目3D人体姿态与形状估计因深度歧义、遮挡和截断而具有固有病态性。现有概率方法通过最大化真实姿态给定图像的似然来学习可能的3D人体网格分布。我们发现仅依赖此目标函数不足以充分捕捉完整分布。为此,我们提出额外用2D姿态检测器热图编码的分布来监督学习到的分布,最小化两者间距离。同时揭示当前方法常为不可见关节生成错误假设,而评估协议未检测此类问题。我们证明训练中使用人体分割掩码可显著减少无效样本,并提出两个新评估指标。基于归一化流的方法生成与图像证据一致且对模糊部位保持高多样性的3D人体网格假设。在3DPW和EMDB数据集上的实验表明,本方法优于其他先进概率方法。代码已公开用于研究目的。
原文摘要 · Abstract (English)
Monocular 3D human pose and shape estimation is an inherently ill-posed problem due to depth ambiguities, occlusions, and truncations. Recent probabilistic approaches learn a distribution over plausible 3D human meshes by maximizing the likelihood of the ground-truth pose given an image. We show that this objective function alone is not sufficient to best capture the full distributions. Instead, we propose to additionally supervise the learned distributions by minimizing the distance to distributions encoded in heatmaps of a 2D pose detector. Moreover, we reveal that current methods often generate incorrect hypotheses for invisible joints which is not detected by the evaluation protocols. We demonstrate that person segmentation masks can be utilized during training to significantly decrease the number of invalid samples and introduce two metrics to evaluate it. Our normalizing flow-based approach predicts plausible 3D human mesh hypotheses that are consistent with the image evidence while maintaining high diversity for ambiguous body parts. Experiments on 3DPW and EMDB show that we outperform other state-of-the-art probabilistic methods. Code is available for research purposes at https://github.com/twehrbein/humr.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。