arXiv:2501.00602cs.CVcs.LG2025-01被引 54

用Transformer一次生成动态室外场景的3D结构,速度快质量高。

STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes

  • 用自监督场景流聚合多帧3D高斯点,实现任意时间视角的完整重建
  • 动态区域重建比顶尖方法提升4.3~6.6 PSNR,大场景重建仅需200ms
  • 无需人工标注即可自动分割动态物体,适合实时渲染与大规模场景理解

我们提出STORM,一种用于从稀疏观测中重建动态室外场景的时空重建模型。现有方法依赖逐场景优化、密集时空观测和强运动监督,导致计算耗时长、泛化能力差,且因噪声伪标签导致质量下降。STORM采用数据驱动的Transformer架构,通过单次前向传播直接推断由3D高斯及其速度参数化的动态3D场景表示。其核心设计是利用自监督场景流聚合所有帧的3D高斯,并变换至目标时刻,实现任意视角在任意时间的完整(即“非遮挡”)重建。作为衍生特性,STORM仅通过重建损失即可自动捕获动态实例并生成高质量掩码。在公开数据集上的实验表明,STORM在动态区域重建精度上超越当前最优逐场景优化方法(+4.3~6.6 PSNR)和现有前馈方法(+2.1~4.7 PSNR),大场景重建仅需200ms,支持实时渲染;在场景流估计方面表现更优,3D EPE降低0.422m,Acc5提升28.02%。此外,我们展示了该模型在四个附加任务中的应用潜力,揭示了自监督学习在动态场景理解中的广泛前景。

原文摘要 · Abstract (English)

We present STORM, a spatio-temporal reconstruction model designed for reconstructing dynamic outdoor scenes from sparse observations. Existing dynamic reconstruction methods often rely on per-scene optimization, dense observations across space and time, and strong motion supervision, resulting in lengthy optimization times, limited generalization to novel views or scenes, and degenerated quality caused by noisy pseudo-labels for dynamics. To address these challenges, STORM leverages a data-driven Transformer architecture that directly infers dynamic 3D scene representations--parameterized by 3D Gaussians and their velocities--in a single forward pass. Our key design is to aggregate 3D Gaussians from all frames using self-supervised scene flows, transforming them to the target timestep to enable complete (i.e., "amodal") reconstructions from arbitrary viewpoints at any moment in time. As an emergent property, STORM automatically captures dynamic instances and generates high-quality masks using only reconstruction losses. Extensive experiments on public datasets show that STORM achieves precise dynamic scene reconstruction, surpassing state-of-the-art per-scene optimization methods (+4.3 to 6.6 PSNR) and existing feed-forward approaches (+2.1 to 4.7 PSNR) in dynamic regions. STORM reconstructs large-scale outdoor scenes in 200ms, supports real-time rendering, and outperforms competitors in scene flow estimation, improving 3D EPE by 0.422m and Acc5 by 28.02%. Beyond reconstruction, we showcase four additional applications of our model, illustrating the potential of self-supervised learning for broader dynamic scene understanding.

3D重建动态场景Transformer实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。