arXiv:2502.03465cs.CVcs.AI2025-02被引 8

用3D高斯流实现单目视频的高效动态建模

Seeing World Dynamics in a Nutshell

  • 将视频视为时空连续的3D高斯流,实现端到端动态重建
  • 单次前向传播完成建模,支持实时应用且保持几何一致性
  • 适合需要真实感动态场景重建的研究者与开发者

我们研究如何以空间-时间一致的方式高效表示随意拍摄的单目视频。现有方法多依赖2D/2.5D技术,将视频视为时空像素集合,因缺乏时间连贯性和显式3D结构,在复杂运动、遮挡和几何一致性方面表现不佳。受单目视频作为动态3D世界投影的启发,我们探索通过时空连续的高斯原型来表示视频的内在3D形式。本文提出NutWorld框架,可在一次前向传播中将单目视频转化为动态3D高斯表示。核心是结构化时空对齐高斯(STAG)表示,实现无需优化的场景建模,并具备有效的深度与光流正则化。大量实验表明,NutWorld在实现高质量视频重建的同时,支持多种下游实时应用。演示与代码将公开于https://github.com/Nut-World/NutWorld。

原文摘要 · Abstract (English)

We consider the problem of efficiently representing casually captured monocular videos in a spatially- and temporally-coherent manner. While existing approaches predominantly rely on 2D/2.5D techniques treating videos as collections of spatiotemporal pixels, they struggle with complex motions, occlusions, and geometric consistency due to absence of temporal coherence and explicit 3D structure. Drawing inspiration from monocular video as a projection of the dynamic 3D world, we explore representing videos in their intrinsic 3D form through continuous flows of Gaussian primitives in space-time. In this paper, we propose NutWorld, a novel framework that efficiently transforms monocular videos into dynamic 3D Gaussian representations in a single forward pass. At its core, NutWorld introduces a structured spatial-temporal aligned Gaussian (STAG) representation, enabling optimization-free scene modeling with effective depth and flow regularization. Through comprehensive experiments, we demonstrate that NutWorld achieves high-fidelity video reconstruction quality while enabling various downstream applications in real-time. Demos and code will be available at https://github.com/Nut-World/NutWorld.

3D重建动态场景高斯渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。