单目视频动态场景重建,速度更快、质量不降。
MOSAIC-GS: Monocular Scene Reconstruction via Advanced Initialization for Complex Dynamic Environments
- 初始化阶段融合深度、光流等多几何线索,预估3D动态
- 动态点用时变傅里叶曲线编码轨迹,支持非刚性变形
- 相比现有方法加速明显,适合实时应用
我们提出MOSAIC-GS,一种基于高斯点云的高效单目动态场景重建新方法。由于单目视角缺乏多视图约束,场景几何与时间一致性恢复极具挑战。为此,我们利用深度、光流、动态物体分割和点追踪等多重几何线索,结合刚性运动约束,在初始化阶段预估初步3D动态。先于光度优化获取动态信息,减少对视觉外观运动推断的依赖,避免歧义。为实现紧凑表示、快速训练与实时渲染,场景被分解为静态与动态部分。动态成分中每个高斯点的运动轨迹由时间依赖的Poly-Fourier曲线表示,实现参数高效编码。实验表明,MOSAIC-GS在标准单目动态场景基准上,相比现有方法显著提升优化与渲染速度,同时保持与顶尖方法相当的重建质量。
原文摘要 · Abstract (English)
We present MOSAIC-GS, a novel, fully explicit, and computationally efficient approach for high-fidelity dynamic scene reconstruction from monocular videos using Gaussian Splatting. Monocular reconstruction is inherently ill-posed due to the lack of sufficient multiview constraints, making accurate recovery of object geometry and temporal coherence particularly challenging. To address this, we leverage multiple geometric cues, such as depth, optical flow, dynamic object segmentation, and point tracking. Combined with rigidity-based motion constraints, these cues allow us to estimate preliminary 3D scene dynamics during an initialization stage. Recovering scene dynamics prior to the photometric optimization reduces reliance on motion inference from visual appearance alone, which is often ambiguous in monocular settings. To enable compact representations, fast training, and real-time rendering while supporting non-rigid deformations, the scene is decomposed into static and dynamic components. Each Gaussian in the dynamic part of the scene is assigned a trajectory represented as time-dependent Poly-Fourier curve for parameter-efficient motion encoding. We demonstrate that MOSAIC-GS achieves substantially faster optimization and rendering compared to existing methods, while maintaining reconstruction quality on par with state-of-the-art approaches across standard monocular dynamic scene benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。