用多尺度动态建模实现单目视频的高保真4D重建
Gaussian Sequences with Multi-Scale Dynamics for 4D Reconstruction from Monocular Casual Videos
- 通过分层运动分解构建多尺度动态表示
- 在基准与真实场景上实现更准确的动态新视角合成
- 适合需要物理合理动态重建的研究者
从日常视频中理解动态场景对可扩展机器人学习至关重要,但严格单目条件下的四维(4D)重建仍高度病态。我们提出的核心见解是:真实世界动态具有从物体到粒子级别的多尺度规律性。为此,设计了多尺度动态机制,将复杂运动场进行因子分解。在此框架下,提出多尺度动态高斯序列,一种通过多层次运动组合生成的动态3D高斯表示。该分层结构显著缓解了重建模糊性,促进物理合理的动态表现。进一步引入视觉基础模型的多模态先验,建立互补监督,约束解空间并提升重建保真度。所提方法实现了单目日常视频的精确且全局一致的4D重建。在基准和真实世界操作数据集上的动态新视角合成实验表明,性能显著优于现有方法。
原文摘要 · Abstract (English)
Understanding dynamic scenes from casual videos is critical for scalable robot learning, yet four-dimensional (4D) reconstruction under strictly monocular settings remains highly ill-posed. To address this challenge, our key insight is that real-world dynamics exhibits a multi-scale regularity from object to particle level. To this end, we design the multi-scale dynamics mechanism that factorizes complex motion fields. Within this formulation, we propose Gaussian sequences with multi-scale dynamics, a novel representation for dynamic 3D Gaussians derived through compositions of multi-level motion. This layered structure substantially alleviates ambiguity of reconstruction and promotes physically plausible dynamics. We further incorporate multi-modal priors from vision foundation models to establish complementary supervision, constraining the solution space and improving the reconstruction fidelity. Our approach enables accurate and globally consistent 4D reconstruction from monocular casual videos. Experiments of dynamic novel-view synthesis (NVS) on benchmark and real-world manipulation datasets demonstrate considerable improvements over existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。