用视觉大模型指导稀疏视角3D重建,提升细节质量
Intern-GS: Vision Model Guided Sparse-View 3D Gaussian Splatting
- 用大模型生成稠密初始点云,替代传统SfM方法
- 在优化中预测未观测视角的深度与外观信息
- 适用于前后向及大规模场景,效果领先
稀疏视角场景重建常因观测数据有限而面临信息不全问题,导致重建质量不佳。为此,我们提出Intern-GS,一种利用视觉基础模型先验知识增强稀疏视角3D高斯点云渲染的新方法。该方法通过视觉基础模型指导3D高斯点云的初始化与优化过程。在初始化阶段,采用DUSt3R生成稠密且无冗余的高斯点云,有效克服传统结构光运动(SfM)方法在稀疏视角下的局限性。在优化阶段,视觉基础模型预测未观测视角的深度与外观,从而修正未覆盖区域的信息缺失。大量实验表明,Intern-GS在多种数据集上均取得最优渲染效果,涵盖前向场景与大规模场景,如LLFF、DTU和Tanks and Temples。
原文摘要 · Abstract (English)
Sparse-view scene reconstruction often faces significant challenges due to the constraints imposed by limited observational data. These limitations result in incomplete information, leading to suboptimal reconstructions using existing methodologies. To address this, we present Intern-GS, a novel approach that effectively leverages rich prior knowledge from vision foundation models to enhance the process of sparse-view Gaussian Splatting, thereby enabling high-quality scene reconstruction. Specifically, Intern-GS utilizes vision foundation models to guide both the initialization and the optimization process of 3D Gaussian splatting, effectively addressing the limitations of sparse inputs. In the initialization process, our method employs DUSt3R to generate a dense and non-redundant gaussian point cloud. This approach significantly alleviates the limitations encountered by traditional structure-from-motion (SfM) methods, which often struggle under sparse-view constraints. During the optimization process, vision foundation models predict depth and appearance for unobserved views, refining the 3D Gaussians to compensate for missing information in unseen regions. Extensive experiments demonstrate that Intern-GS achieves state-of-the-art rendering quality across diverse datasets, including both forward-facing and large-scale scenes, such as LLFF, DTU, and Tanks and Temples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。