arXiv:2412.19089cs.CV2024-12ICCV被引 7

用人体姿态作为校准基准,实现无同步无标定视频的动态3D重建。

Humans as a Calibration Pattern: Dynamic 3D Scene Reconstruction from Unsynchronized and Uncalibrated Videos

  • 利用人体姿态估计提供初始约束,解决多视角视频不同步问题。
  • 在真实复杂场景中实现高精度时空校准与高质量动态3D重建。
  • 适合缺乏相机标定信息、仅含人体运动的视频重建任务。

现有动态3D神经场重建方法通常要求输入为同步多视角视频且相机位姿已知,但在真实场景中常无法满足。本文提出,只要视频中包含人体运动,即可通过人体形状和姿态估计来实现动态神经场重建。尽管估计存在噪声,但其可作为高非凸、欠约束优化问题的良好初始化。基于单帧人体形状与姿态参数,我们推导出视频间时间偏移量,并据此估算相机位姿,分析3D关节位置。随后,采用多分辨率网格训练动态神经场,同时协同优化时间偏移与相机位姿。为应对大量待优化参数带来的不稳定性,引入鲁棒的渐进式学习策略。实验表明,该方法在复杂条件下实现了精确的时空校准与高质量场景重建。

原文摘要 · Abstract (English)

Recent works on dynamic 3D neural field reconstruction assume the input from synchronized multi-view videos whose poses are known. The input constraints are often not satisfied in real-world setups, making the approach impractical. We show that unsynchronized videos from unknown poses can generate dynamic neural fields as long as the videos capture human motion. Humans are one of the most common dynamic subjects captured in videos, and their shapes and poses can be estimated using state-of-the-art libraries. While noisy, the estimated human shape and pose parameters provide a decent initialization point to start the highly non-convex and under-constrained problem of training a consistent dynamic neural representation. Given the shape and pose parameters of humans in individual frames, we formulate methods to calculate the time offsets between videos, followed by camera pose estimations that analyze the 3D joint positions. Then, we train the dynamic neural fields employing multiresolution grids while we concurrently refine both time offsets and camera poses. The setup still involves optimizing many parameters; therefore, we introduce a robust progressive learning strategy to stabilize the process. Experiments show that our approach achieves accurate spatio-temporal calibration and high-quality scene reconstruction in challenging conditions.

3D重建动态神经场人体姿态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。