arXiv:2607.03209cs.CVcs.GR2026-07中稿 · ICECET 2026

用基础模型快速初始化3D高斯点云,三分钟完成高质量重建

Fast 3D Foundation Model Initialized Gaussian Splatting

论文配图:Fast 3D Foundation Model Initialized Gaussian Splatting
图 1 · 摘自论文原文
  • 用3D基础模型直接初始化相机位姿和点云,跳过传统SfM流程
  • 仅需50-60张图即可收敛,重建质量达23.61 dB PSNR、0.19 LPIPS
  • 适合机器人、VR等需要近实时3D重建的应用场景

本文提出一种无需传统运动恢复结构(SfM)的快速高保真3D高斯点云渲染方法。该方法利用3D基础模型(3DFM)进行相机位姿与点云的初始估计,并通过深度引导损失函数联合优化相机位姿与高斯原语。即使在粗略初始化下,仅需50-60张输入图像即可实现快速收敛。为提升稀疏视角下的重建质量,引入基于MLP的位姿精修模块,并结合基础模型提供的深度监督。在Mip-NeRF 360、Tanks and Temples和RobustNeRF数据集上的实验表明,该方法在约三分钟内完成每场景训练,达到23.61 dB PSNR和0.19 LPIPS的竞争力性能。所提方法可快速生成可直接使用的3DGS模型,适用于机器人、虚拟现实及自动驾驶等近实时应用。

原文摘要 · Abstract (English)

This paper introduces a fast method for high-quality 3D Gaussian Splatting (3DGS) reconstruction without traditional Structure-from-Motion (SfM). The proposed approach leverages 3D Foundation Models (3DFMs) for camera pose and point-cloud initialization, then jointly optimizes both camera poses and Gaussian primitives using a depth-guided loss function. This enables fast convergence even from rough initialization with as few as 50-60 input views. To further improve reconstruction quality in sparse-view scenarios, an MLP-based pose refinement module is introduced alongside depth-guided supervision from the foundation model. Extensive experiments on Mip-NeRF 360, Tanks and Temples, and RobustNeRF demonstrate that the proposed method achieves competitive reconstruction quality (23.61 dB PSNR, 0.19 LPIPS) while reducing training time to approximately three minutes per scene. The proposed method produces ready-to-use 3DGS models at a fraction of the time required by existing pipelines, making it suitable for near real-time applications in robotics, VR, and autonomous navigation.

3D重建高斯溅射基础模型实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。