从单图精准恢复人体3D网格,解决近景拍摄时的投影失真问题。
BLADE: Single-view Body Mesh Learning through Accurate Depth Estimation
- 基于深度与透视畸变的反向关系,直接估计人物前后位置Tz
- 首次实现单图下透视参数的准确恢复,2D对齐与3D姿态精度双提升
- 适用于近景图像,适合动作捕捉、虚拟试衣等实际场景
单图人体网格重建因同时需估计身体形状、姿态和相机参数而具有病态性。现有方法在远距离图像上表现良好,但在人物靠近镜头时失效。当前方法难以兼顾3D姿态与2D对齐精度,主要源于由正交参数推导的启发式透视投影误差。为解决这一长期挑战,我们提出BLADE方法,无需启发假设即可从单图准确恢复透视参数。我们发现透视畸变与人物Z向位移Tz存在逆相关,且可通过图像可靠估计Tz。进一步表明,准确估计Tz对近景图像中的人体网格重建至关重要。一旦获得Tz与3D人体网格,即可精确恢复焦距与完整3D平移。在标准基准与真实近距图像上的大量实验表明,本方法是首个能从单图准确恢复投影参数的方法,从而在多种图像上实现3D姿态估计与2D对齐的最先进性能。
原文摘要 · Abstract (English)
Single-image human mesh recovery is a challenging task due to the ill-posed nature of simultaneous body shape, pose, and camera estimation. Existing estimators work well on images taken from afar, but they break down as the person moves close to the camera. Moreover, current methods fail to achieve both accurate 3D pose and 2D alignment at the same time. Error is mainly introduced by inaccurate perspective projection heuristically derived from orthographic parameters. To resolve this long-standing challenge, we present our method BLADE which accurately recovers perspective parameters from a single image without heuristic assumptions. We start from the inverse relationship between perspective distortion and the person's Z-translation Tz, and we show that Tz can be reliably estimated from the image. We then discuss the important role of Tz for accurate human mesh recovery estimated from close-range images. Finally, we show that, once Tz and the 3D human mesh are estimated, one can accurately recover the focal length and full 3D translation. Extensive experiments on standard benchmarks and real-world close-range images show that our method is the first to accurately recover projection parameters from a single image, and consequently attain state-of-the-art accuracy on 3D pose estimation and 2D alignment for a wide range of images. https://research.nvidia.com/labs/amri/projects/blade/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。