无需优化即可将3DGS重建转为结构化潜空间,支持大规模航拍场景生成。
GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation

- 将预优化的3DGS直接转为稀疏体素,不需额外优化
- 潜空间大小随占据体素数增长,可支持海量点云
- 适用于卫星图像驱动的大范围航拍场景生成
许多可扩展的隐式3D生成器基于结构化张量,而预优化的3D高斯溅射(3DGS)重建是无序、空间不规则且原始数量差异大。本文提出GS-Voxel,一种无需拟合的结构化潜空间框架,用于大规模航拍3DGS场景生成。该方法确定性地将兼容的预优化3DGS重建转换为稀疏活跃体素,无需额外每场景优化,保留选定原语的亚体素位置和渲染属性。随后,通过专用的分解变分自编码器(VAE)分别将体素几何与局部高斯属性编码为稀疏3D潜变量,其大小随占据体素数增长,而非受固定场景级原始数限制。我们在GS-Voxel潜空间中训练图像条件流模型以生成航拍3DGS场景。关键应用之一是大区域场景生成:重叠感知的分块推理使合成突破单个训练作物的限制,基于卫星视图图像进行扩展。结果表明,GS-Voxel为预优化的航拍3DGS重建提供了可扩展的结构化潜空间,潜容量随占据体素数动态增长。
原文摘要 · Abstract (English)
Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) reconstructions are unordered, spatially irregular, and vary widely in primitive count. We present GS-Voxel, a fitting-free structured latent framework, and evaluate it for large-scale aerial 3D Gaussian scene generation. GS-Voxel deterministically converts a compatible pre-optimized 3DGS reconstruction into sparse active voxels without additional per-scene optimization, retaining the sub-voxel positions and rendering attributes of the selected primitives. A GS-specific factorized VAE then separately encodes voxel geometry and local Gaussian attributes into sparse 3D latents whose size grows with the number of occupied voxels rather than being limited by a fixed scene-wide primitive count. We train image-conditioned flow models in the GS-Voxel latent space to generate aerial 3DGS scenes. A key application enabled by GS-Voxel is large-area scene generation: overlap-aware tiled inference extends synthesis beyond a single training crop conditioned on satellite-view images. Our results show that GS-Voxel provides structured latents for pre-optimized aerial 3DGS reconstructions, with latent capacity that grows with the number of occupied voxels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。