arXiv:2506.05558cs.CV2025-06被引 43

实时重建大场景新视角图像,捕获即生成高质量3D高斯点云。

On-the-fly Reconstruction for Large-Scale Novel View Synthesis from Unposed Images

  • 捕获时同步估计相机位姿并增量生成3DGS,实现边拍边建
  • 通过快速初始位姿+增量点云采样,训练速度提升数倍
  • 支持大场景、宽基线拍摄,适合实地快速三维重建

基于辐射场的方法如3D高斯泼溅(3DGS)可从照片轻松重建,实现自由视角导航。然而,使用运动结构法(SfM)和3DGS优化进行位姿估计仍需数分钟至数小时。结合SLAM与3DGS的方法虽快,但难以处理宽基线和大场景。本文提出一种在拍摄后立即生成相机位姿和训练好的3DGS的实时方法。该方法可处理密集、宽基线的有序照片序列及大尺度场景。首先引入快速初始位姿估计,利用学习特征和适合GPU的微型束调整;随后直接采样高斯原始位置与形状,按需增量生成,显著加速训练。这两项高效步骤使位姿与高斯原始参数能快速稳健联合优化。通过渐进式聚类3DGS原始点、存入锚点并卸载至显存外,实现可扩展的辐射场构建。聚类后的原始点逐步合并,保持任意视角下的3DGS规模可控。我们在多种数据集上评估,结果表明本方法能实时处理目标所有拍摄场景与场景大小,在速度、图像质量上均与仅针对特定风格或尺寸的方法竞争相当。

原文摘要 · Abstract (English)

Radiance field methods such as 3D Gaussian Splatting (3DGS) allow easy reconstruction from photos, enabling free-viewpoint navigation. Nonetheless, pose estimation using Structure from Motion and 3DGS optimization can still each take between minutes and hours of computation after capture is complete. SLAM methods combined with 3DGS are fast but struggle with wide camera baselines and large scenes. We present an on-the-fly method to produce camera poses and a trained 3DGS immediately after capture. Our method can handle dense and wide-baseline captures of ordered photo sequences and large-scale scenes. To do this, we first introduce fast initial pose estimation, exploiting learned features and a GPU-friendly mini bundle adjustment. We then introduce direct sampling of Gaussian primitive positions and shapes, incrementally spawning primitives where required, significantly accelerating training. These two efficient steps allow fast and robust joint optimization of poses and Gaussian primitives. Our incremental approach handles large-scale scenes by introducing scalable radiance field construction, progressively clustering 3DGS primitives, storing them in anchors, and offloading them from the GPU. Clustered primitives are progressively merged, keeping the required scale of 3DGS at any viewpoint. We evaluate our solution on a variety of datasets and show that our solution can provide on-the-fly processing of all the capture scenarios and scene sizes we target while remaining competitive with other methods that only handle specific capture styles or scene sizes in speed, image quality, or both.

3D重建3DGS实时渲染新视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。