用卫星图生成无缝城市级点云,支持大规模真实场景建模。
GridFlow: Structured Latent Flow for Seamless City-Scale 3D Point Cloud Generation

- 分阶段构建网格对齐的隐空间流,实现跨区块无缝衔接。
- 每150×150米区块生成10万点,几何与颜色分别优化。
- 提出新基准数据集,支持城市级点云生成评估。
从遥感数据生成逼真的城市三维环境对于仿真、城市规划和混合现实至关重要,但现有点云生成方法仅限于单个物体或封闭室内场景,难以应对城市尺度下的规模、无缝拼接和部分可观测性挑战。我们提出 extit{GridFlow},一种多阶段框架,可基于卫星影像、语义分割图和数字地表模型(DSM),生成稠密彩色点云(每150m×150m区块含10⁵点)。其核心为网格对齐的变分自编码器(Grid-Aligned VAE),将每个区块编码为拓扑保持的隐空间网格,使隐变量对应固定空间区域,从而实现空间一致的多模态条件输入,并在隐空间中隐式对齐数千个边界点,保障跨区块生成的无缝性。随后,条件化修正流模型从融合的多模态输入中合成几何隐变量,而面向朝向的扩散着色器则分别处理可见水平面与遮挡立面。为支持标准化评估,我们基于公开3D数据源构建了 extit{City3D-MultiGen} 基准,包含来自墨尔本与伦敦的16.3万张密集标注区块,涵盖对齐的点云、卫星图像、语义地图与高程数据。实验表明, extit{GridFlow} 在所有几何指标上均优于适配的点云生成基线,生成的彩色点云视觉连贯,且可在任意大范围城市区域内实现无缝边界。
原文摘要 · Abstract (English)
Generating realistic 3D city environments from remote sensing data is important for simulation, urban planning, and mixed reality, yet existing point cloud generation methods are limited to single objects or bounded indoor scenes and cannot handle the scale, seamless tiling, and partial observability challenges of city-scale generation. We present \ours{}, a multi-stage framework that generates dense, colored point clouds ($10^5$ points per $150\text{m}{\times}150\text{m}$ tile) at city scale, conditioned on satellite imagery, semantic segmentation maps, and digital surface models (DSM). A \emph{Grid-Aligned VAE} encodes each tile into a topology-preserving latent grid where tokens correspond to fixed spatial regions, enabling spatially coherent multi-modal conditioning and compact latent-space edge consistency that implicitly aligns thousands of boundary points for seamless cross-tile generation. A conditional rectified flow model synthesizes geometry latents from the fused multi-modal conditions, and an orientation-aware diffusion colorizer separately handles satellite-visible horizontal surfaces and occluded vertical façades. To support standardized evaluation, we build on public 3D data sources to introduce \emph{City3D-MultiGen}, a benchmark of $163$K densely annotated tiles from Melbourne and London with aligned point clouds, satellite images, semantic maps, and elevation data. Experiments show that \ours{} outperforms adapted point cloud generation baselines across all geometry metrics and produces visually coherent colored point clouds with seamless boundaries over arbitrarily large urban extents. Our benchmark details are available at https://huggingface.co/datasets/e32/City3D-MultiGen
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。