无需标定相机,一步恢复多人动作与视角参数。
Simultaneously Recovering Multi-Person Meshes and Multi-View Cameras with Human Semantics
- 利用人体语义初始化相机参数,免去传统标定流程。
- 通过姿态-几何一致性关联多视角人体,提升重建精度。
- 引入潜在运动先验,实现相机与动作联合优化,适合动态多人场景。
动态多人网格恢复在体育转播、虚拟现实和视频游戏中有广泛应用。然而,现有多视角框架依赖耗时的相机标定过程。本文聚焦于未标定相机下的多人动作捕捉,主要面临两大挑战:人与人之间的交互和遮挡导致相机标定与动作捕捉固有模糊;缺乏密集对应关系难以约束动态场景中的稀疏相机几何结构。核心思路是结合运动先验知识,从噪声人体语义中同时估计相机参数与人体网格。首先利用2D图像中的人体信息初始化相机内参与外参,不依赖任何额外标定工具或背景特征。其次引入姿态-几何一致性,关联不同视角检测到的人体。最后提出潜在运动先验,优化相机参数与人体运动。实验表明,可通过一步重建获得精确的相机参数与人体动作。代码已公开于https://github.com/boycehbz/DMMR。
原文摘要 · Abstract (English)
Dynamic multi-person mesh recovery has broad applications in sports broadcasting, virtual reality, and video games. However, current multi-view frameworks rely on a time-consuming camera calibration procedure. In this work, we focus on multi-person motion capture with uncalibrated cameras, which mainly faces two challenges: one is that inter-person interactions and occlusions introduce inherent ambiguities for both camera calibration and motion capture; the other is that a lack of dense correspondences can be used to constrain sparse camera geometries in a dynamic multi-person scene. Our key idea is to incorporate motion prior knowledge to simultaneously estimate camera parameters and human meshes from noisy human semantics. We first utilize human information from 2D images to initialize intrinsic and extrinsic parameters. Thus, the approach does not rely on any other calibration tools or background features. Then, a pose-geometry consistency is introduced to associate the detected humans from different views. Finally, a latent motion prior is proposed to refine the camera parameters and human motions. Experimental results show that accurate camera parameters and human motions can be obtained through a one-step reconstruction. The code are publicly available at~\url{https://github.com/boycehbz/DMMR}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。