从单目视频生成可自由渲染的动态3D高斯模型
World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video

- 用像素对齐的渲染图条件化视频模型,融合外观、几何与运动信息
- 在真实视频上实现4D重建新纪录,支持大视角变化和动态场景
- 适合做动态3D重建、视觉编辑或虚拟现实内容生成的研究者
我们提出World from Motion,一种从单目视频生成可自由渲染的动态3D高斯表示的方法。该方法将视频模型以密集像素对齐的渲染结果为条件,这些渲染结果编码了沿输入和目标相机轨迹的外观、几何及3D场景运动,用于修正渲染伪影并填补初始重建中的缺失区域。为训练此模型,我们构建了一个包含多视角视频对与动态3DGS表示的对齐数据集,其中包含单目重建特有的模拟伪影。测试时,我们将模型生成的内容(包括新观测区域和运动)提炼回单一一致且高质量的动态3DGS,同时提升新视角合成与底层3D运动的准确性。该方法在4D重建任务上达到新基准,并能无缝推广至具有大幅视角变化和动态运动的真实视频。
原文摘要 · Abstract (English)
We present World from Motion, a method for generating freely renderable dynamic 3D Gaussian representations from monocular videos. Our approach conditions a video model on dense, pixel-aligned renderings that encode appearance, geometry, and 3D scene motion along both input and target camera trajectories to correct rendering artifacts and fill in missing regions from an initial reconstruction. To train this model, we construct a dataset of aligned multiview video pairs and dynamic 3DGS representations, with simulated artifacts characteristic of monocular reconstruction. At test time, we distill the model's generations, including newly observed regions and motions, back into a single consistent, high-quality dynamic 3DGS, improving both novel-view synthesis and the underlying 3D motion. Our method sets a new state of the art in 4D reconstruction and seamlessly generalizes to in-the-wild videos with large viewpoint changes and dynamic motions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。