单目视频中一次性重建动态人和静态场景的高质高效方法
GUSH3R: Everyone Everywhere All at Once as Gaussians

- 用3D高斯点云统一表示人与场景,单次前向传播完成重建
- 在多个数据集上实现媲美优化法的视图合成质量,推理速度更快
- 适合需要实时动态人-场景重建的应用,如VR/AR和数字人
从单目视频重建动态人-场景环境是一项挑战性任务,需联合建模场景几何、相机运动及非刚性人体动态,并支持逼真渲染。现有前馈方法虽能高效预测几何,但通常仅限于点云、网格等非逼真表示,或无法处理动态人体。为此,我们提出GUSH3R(Gaussian-Unified Scene Human 3D Reconstruction),一种用于在线动态人-场景重建的前馈框架。从单目人-场景视频中,该方法以3D高斯溅射(3DGS)基元形式,一次性重建动态人(everyone)和静态场景(everywhere),实现几何一致且支持新视角合成。在单目人-场景数据集上的实验表明,本方法在新视角合成质量上达到竞争水平,同时相比基于优化的方法显著提升推理效率。
原文摘要 · Abstract (English)
Reconstructing dynamic human-scene environments from monocular videos is a challenging problem that requires jointly modeling scene geometry, camera motion, and non-rigid human dynamics while enabling photorealistic rendering. Recent feed-forward methods can efficiently predict geometry, but they are often limited to non-photorealistic representations such as point clouds and meshes, or they fail to handle non-rigid objects, particularly dynamic humans. To fill this gap, we present GUSH3R (Gaussian-Unified Scene Human 3D Reconstruction), a feed-forward framework for online dynamic human-scene reconstruction. From a monocular human-scene video, our method reconstructs dynamic humans (everyone) and static scenes (everywhere) in a single forward pass (all at once) as 3D Gaussian Splatting (3DGS) primitives (as gaussians), which are geometrically consistent and capable of novel view synthesis. Experiments on monocular human-scene datasets demonstrate that our approach achieves competitive novel view synthesis quality while significantly improving inference efficiency compared to optimization-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。