无需环境线索,自动校准相机角度实现精准人体三维重建
From Camera to World: A Plug-and-Play Module for Human Mesh Transformation
- 用人体空间结构预测相机俯仰角,避免依赖外部环境信息
- 在SPEC-SYN和SPEC-MTP数据集上优于现有最佳方法
- 可直接插入现有流程,适合需高精度人体姿态重建的场景
从真实场景图像中重建世界坐标系下的精确3D人体网格仍具挑战性,因缺乏相机旋转信息。现有方法虽在假设相机无旋转时取得良好效果,但转换至世界坐标系后会产生显著误差。为此,我们提出Mesh-Plug——一种即插即用模块,可准确将人体网格从相机坐标系转换到世界坐标系。核心创新在于以人为中心的方法:利用原始网格渲染的深度图与图像,联合估计相机旋转参数,无需依赖环境线索。具体而言,先训练一个聚焦人体空间构型的相机俯仰角预测模块;再结合预测参数,设计网格调整模块,同步优化根关节朝向与身体姿态。大量实验表明,该框架在基准数据集SPEC-SYN和SPEC-MTP上均超越当前最优方法。
原文摘要 · Abstract (English)
Reconstructing accurate 3D human meshes in the world coordinate system from in-the-wild images remains challenging due to the lack of camera rotation information. While existing methods achieve promising results in the camera coordinate system by assuming zero camera rotation, this simplification leads to significant errors when transforming the reconstructed mesh to the world coordinate system. To address this challenge, we propose Mesh-Plug, a plug-and-play module that accurately transforms human meshes from camera coordinates to world coordinates. Our key innovation lies in a human-centered approach that leverages both RGB images and depth maps rendered from the initial mesh to estimate camera rotation parameters, eliminating the dependency on environmental cues. Specifically, we first train a camera rotation prediction module that focuses on the human body's spatial configuration to estimate camera pitch angle. Then, by integrating the predicted camera parameters with the initial mesh, we design a mesh adjustment module that simultaneously refines the root joint orientation and body pose. Extensive experiments demonstrate that our framework outperforms state-of-the-art methods on the benchmark datasets SPEC-SYN and SPEC-MTP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。