仅用两张无姿态图像实现动态场景高保真3D重建
DynSUP: Dynamic Gaussian Splatting from An Unposed Image Pair
- 通过物体级两视图捆绑调整分解动态场景为刚性块,联合估计相机与物体运动
- 引入SE(3)场驱动的高斯训练,实现每个高斯的细粒度运动建模
- 适用于稀疏视角、未知姿态的动态场景,适合视频生成与虚拟拍摄应用
最近的3D高斯点云渲染取得了显著进展。现有方法通常假设场景静态或需要多张带已知位姿的图像。动态性、视角稀疏和未知位姿会因几何约束不足而大幅增加问题复杂度。为此,我们提出一种仅需两张无位姿图像即可在动态环境中拟合高斯的方法。首先,提出物体级两视图捆绑调整策略,将动态场景分解为分段刚性组件,并联合估计相机位姿与动态物体运动。其次,设计基于SE(3)场驱动的高斯训练方法,通过可学习的逐高斯变换实现精细运动建模。本方法在保持时间一致性的同时,实现了高质量的新视角合成。在合成与真实数据集上的实验表明,该方法显著优于针对静态场景、多图像或已知位姿设计的最先进方法。
原文摘要 · Abstract (English)
Recent advances in 3D Gaussian Splatting have shown promising results. Existing methods typically assume static scenes and/or multiple images with prior poses. Dynamics, sparse views, and unknown poses significantly increase the problem complexity due to insufficient geometric constraints. To overcome this challenge, we propose a method that can use only two images without prior poses to fit Gaussians in dynamic environments. To achieve this, we introduce two technical contributions. First, we propose an object-level two-view bundle adjustment. This strategy decomposes dynamic scenes into piece-wise rigid components, and jointly estimates the camera pose and motions of dynamic objects. Second, we design an SE(3) field-driven Gaussian training method. It enables fine-grained motion modeling through learnable per-Gaussian transformations. Our method leads to high-fidelity novel view synthesis of dynamic scenes while accurately preserving temporal consistency and object motion. Experiments on both synthetic and real-world datasets demonstrate that our method significantly outperforms state-of-the-art approaches designed for the cases of static environments, multiple images, and/or known poses. Our project page is available at https://colin-de.github.io/DynSUP/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。