单目视频实时重建人与场景,无需预标定且速度快。
ODHSR: Online Dense 3D Reconstruction of Humans and Scenes from Monocular Videos
- 在线联合追踪相机、姿态并重建3D人体与场景。
- 用3D高斯点云实现高效建模,支持新视角合成。
- 适合需要快速生成逼真人-场景3D内容的开发者。
从单目野外视频中构建逼真的三维人与场景,是理解以人为中心的三维世界的关键。现有神经渲染方法虽能实现整体重建,但需预先标定相机和人体姿态,且训练耗时长达数天。本文提出一种新型统一框架,首次实现相机追踪、人体姿态估计与人体-场景重建的在线联合处理。采用3D高斯点云(3D Gaussian Splatting)高效学习人体与场景的高斯基元;设计基于重建的相机追踪与姿态估计模块,实现姿态与外观的有效解耦。特别地,引入人体形变模块以还原细节并提升对分布外姿态的泛化能力;为准确建模人体与场景的空间关联,提出遮挡感知的人体轮廓渲染与单目几何先验,显著提升重建质量。在EMDB与NeuMan数据集上的实验表明,该方法在相机追踪、姿态估计、新视角合成及运行效率方面均达到或超越现有水平。
原文摘要 · Abstract (English)
Creating a photorealistic scene and human reconstruction from a single monocular in-the-wild video figures prominently in the perception of a human-centric 3D world. Recent neural rendering advances have enabled holistic human-scene reconstruction but require pre-calibrated camera and human poses, and days of training time. In this work, we introduce a novel unified framework that simultaneously performs camera tracking, human pose estimation and human-scene reconstruction in an online fashion. 3D Gaussian Splatting is utilized to learn Gaussian primitives for humans and scenes efficiently, and reconstruction-based camera tracking and human pose estimation modules are designed to enable holistic understanding and effective disentanglement of pose and appearance. Specifically, we design a human deformation module to reconstruct the details and enhance generalizability to out-of-distribution poses faithfully. Aiming to learn the spatial correlation between human and scene accurately, we introduce occlusion-aware human silhouette rendering and monocular geometric priors, which further improve reconstruction quality. Experiments on the EMDB and NeuMan datasets demonstrate superior or on-par performance with existing methods in camera tracking, human pose estimation, novel view synthesis and runtime. Our project page is at https://eth-ait.github.io/ODHSR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。