无需深度模型,30秒内快速重建单目视频的3D高精度场景。
KeyGS: A Keyframe-Centric Gaussian Splatting Method for Monocular Image Sequences
- 先用SfM快速获取粗略相机位姿,再通过3D高斯点云优化位姿
- 训练时间从数小时缩短至几分钟,新视角合成更准确
- 分频级联密度化策略,避免位姿漂移,适合实时3D重建
从稀疏2D图像重建高质量3D模型在计算机视觉中备受关注。近期,3D高斯溅射(3DGS)因其显式表示、高效训练速度和实时渲染能力受到重视。然而,现有方法仍严重依赖精确的相机位姿进行重建。尽管部分最新方法尝试在无结构光流(SfM)预处理的单目视频数据集上训练3DGS模型,但这些方法训练时间过长,难以应用于实际场景。本文提出一种无需任何深度或匹配模型的高效框架。该方法首先利用SfM在数秒内快速获取粗略相机位姿,随后借助3DGS中的密集表示对位姿进行精细化优化。此外,我们将密度化过程与联合优化结合,并提出一种由粗到精的频率感知密度化策略,以重建不同层次的细节。该方法有效防止了因高频信号导致的位姿估计陷入局部最优或漂移问题。实验表明,本方法将训练时间从数小时大幅缩短至数分钟,同时在新视角合成和相机位姿估计方面均优于先前方法。
原文摘要 · Abstract (English)
Reconstructing high-quality 3D models from sparse 2D images has garnered significant attention in computer vision. Recently, 3D Gaussian Splatting (3DGS) has gained prominence due to its explicit representation with efficient training speed and real-time rendering capabilities. However, existing methods still heavily depend on accurate camera poses for reconstruction. Although some recent approaches attempt to train 3DGS models without the Structure-from-Motion (SfM) preprocessing from monocular video datasets, these methods suffer from prolonged training times, making them impractical for many applications. In this paper, we present an efficient framework that operates without any depth or matching model. Our approach initially uses SfM to quickly obtain rough camera poses within seconds, and then refines these poses by leveraging the dense representation in 3DGS. This framework effectively addresses the issue of long training times. Additionally, we integrate the densification process with joint refinement and propose a coarse-to-fine frequency-aware densification to reconstruct different levels of details. This approach prevents camera pose estimation from being trapped in local minima or drifting due to high-frequency signals. Our method significantly reduces training time from hours to minutes while achieving more accurate novel view synthesis and camera pose estimation compared to previous methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。