arXiv:2410.20030cs.CVcs.AI2024-10NeurIPS被引 61

用少量图像快速重建大规模精细3D场景,生成结果清晰锐利。

SCube: Instant Large-Scale Scene Reconstruction using VoxSplats

  • 用高分辨率稀疏体素支撑的3D高斯点集表示场景
  • 仅需3张非重叠图像,20秒内生成百万级高斯点,覆盖数百米
  • 适合需要快速高质量3D重建的自动驾驶与生成应用

我们提出SCube,一种从稀疏姿态图像中重建大规模3D场景(几何、外观和语义)的新方法。该方法使用新型表示VoxSplat——基于高分辨率稀疏体素网格的3D高斯点集。通过条件于输入图像的分层体素潜空间扩散模型,逐步生成高分辨率网格,并结合前馈外观预测模型在每个体素内生成高斯点。仅需3张非重叠输入图像,即可在20秒内完成1024³体素网格的重建,生成数百万个高斯点,覆盖数百米范围。现有方法或依赖每场景优化导致视野外无法重建(需密集视角),或依赖低分辨率几何先验导致模糊输出。SCube采用高分辨率稀疏网络,实现少视图下的清晰重建。我们在Waymo自驾车数据集上验证其优越性,并展示了其在LiDAR模拟与文本到场景生成中的应用。

原文摘要 · Abstract (English)

We present SCube, a novel method for reconstructing large-scale 3D scenes (geometry, appearance, and semantics) from a sparse set of posed images. Our method encodes reconstructed scenes using a novel representation VoxSplat, which is a set of 3D Gaussians supported on a high-resolution sparse-voxel scaffold. To reconstruct a VoxSplat from images, we employ a hierarchical voxel latent diffusion model conditioned on the input images followed by a feedforward appearance prediction model. The diffusion model generates high-resolution grids progressively in a coarse-to-fine manner, and the appearance network predicts a set of Gaussians within each voxel. From as few as 3 non-overlapping input images, SCube can generate millions of Gaussians with a 1024^3 voxel grid spanning hundreds of meters in 20 seconds. Past works tackling scene reconstruction from images either rely on per-scene optimization and fail to reconstruct the scene away from input views (thus requiring dense view coverage as input) or leverage geometric priors based on low-resolution models, which produce blurry results. In contrast, SCube leverages high-resolution sparse networks and produces sharp outputs from few views. We show the superiority of SCube compared to prior art using the Waymo self-driving dataset on 3D reconstruction and demonstrate its applications, such as LiDAR simulation and text-to-scene generation.

3D重建高斯点稀疏体素快速重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。