一秒钟从单目视频重建动态4D场景,支持多任务零样本应用。
MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second
- 用像素对齐的高斯原语建模动态3D场景,显式监督运动变化。
- 统一处理外观、几何与运动,实现端到端的重建与视图合成。
- 支持零样本的光流估计和动对象分割,适合实时动态场景应用。
我们提出MoVieS,一种面向单目视频的运动感知视图合成模型,可在1秒内重建4D动态场景。该模型采用像素对齐的高斯原语表示动态3D场景,并显式监督其随时间变化的运动特性。这首次实现了从单目视频中统一建模外观、几何与运动,并在单一学习框架内完成重建、视图合成与3D点追踪。通过将视图合成与几何重建相融合,MoVieS可在多样数据集上进行大规模训练,几乎无需特定任务的监督。因此,它自然支持多种零样本应用,如场景光流估计与运动物体分割。大量实验验证了MoVieS在多任务上的有效性与高效性,在保持竞争力表现的同时,实现数个数量级的速度提升。
原文摘要 · Abstract (English)
We present MoVieS, a Motion-aware View Synthesis model that reconstructs 4D dynamic scenes from monocular videos in one second. It represents dynamic 3D scenes with pixel-aligned Gaussian primitives and explicitly supervises their time-varying motions. This allows, for the first time, the unified modeling of appearance, geometry and motion from monocular videos, and enables reconstruction, view synthesis and 3D point tracking within a single learning-based framework. By bridging view synthesis with geometry reconstruction, MoVieS enables large-scale training on diverse datasets with minimal dependence on task-specific supervision. As a result, it also naturally supports a wide range of zero-shot applications, such as scene flow estimation and moving object segmentation. Extensive experiments validate the effectiveness and efficiency of MoVieS across multiple tasks, achieving competitive performance while offering several orders of magnitude speedups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。