用一张卫星图生成可控制外观的3D城市建筑网格,支持数字孪生与城市仿真。
Sat2City v2: Native 3D City Asset Generation from a Single Satellite Image

- 将真实卫星图与预训练3D隐空间结合,通过几何流和纹理锚定生成精细网格。
- 在24个区域16,241对数据上实现高精度地形重建与外观可控生成。
- 首个公开的卫星-网格配对数据集,适合城市建模与地理智能研究者使用。
从单张卫星图像生成显式3D城市资产对数字孪生、城市模拟和地理空间智能至关重要。与卫星到街景合成不同,该任务需要可重用的带纹理网格,具备合理几何结构和可控外观,而非仅优化渲染少数图像或视频的3D代理。ICCV Sat2City框架首次尝试以卫星导出的高度图为条件,使用级联稀疏体素隐扩散模型生成,但其外观随机,训练数据为合成数据,且特定任务的VAE难以扩展至噪声真实的重建。我们提出Sat2City v2,一种期刊扩展版本,将预训练的原生结构化隐空间3D基础模型适配至弱对齐的卫星图像与带纹理网格。构建了覆盖9个城市24个区域的16,241对真实卫星-网格数据集。不同于从噪声城市网格学习3D表示,Sat2City v2将每个网格编码进预训练的原生3D隐空间,微调卫星条件下的几何流,并利用解码形状锚定卫星条件纹理。该方法保留几何-外观级联架构,同时实现从卫星输入的外观可控生成。在米级DSM重建及几何与外观生成基准测试中,Sat2City v2优于所有对比基线。整体而言,该工作推动卫星到城市生成从渲染导向的3D代理转向显式带纹理网格资产,基于目前已知首个为该资产级任务收集的匹配地理裁剪卫星-网格配对数据集。项目页:https://ai4city-hkust.github.io/Sat2City-v2/
原文摘要 · Abstract (English)
Generating explicit 3D city assets from a single satellite image is important for digital twins, urban simulation, and geospatial intelligence. Unlike satellite-to-street-view synthesis, the task requires a reusable textured mesh with plausible geometry and controllable appearance rather than a 3D proxy optimized only for rendering a small set of images or videos. The ICCV Sat2City framework made a first step by conditioning cascaded sparse-voxel latent diffusion on satellite-derived height maps, but its appearance was random, its training data were synthetic, and its task-specific VAE did not scale well to noisy real-world reconstructions. We present Sat2City v2, a journal extension that adapts a pretrained native structured-latent 3D foundation model to weakly aligned satellite images and textured meshes. We build a real-world dataset with 16,241 satellite-mesh pairs across 24 regions in 9 cities. Instead of learning a 3D representation from noisy city meshes, Sat2City v2 encodes each mesh into a pretrained native 3D latent space, fine-tunes a satellite-conditioned geometry flow, and uses the decoded shape to anchor satellite-conditioned texturing. This retains Sat2City's geometry-to-appearance cascade while enabling appearance-controllable generation from the satellite input. Experiments on metric-scale DSM reconstruction and generative city-asset benchmarks for geometry and appearance show that Sat2City v2 achieves the best overall performance among evaluated baselines. Overall, Sat2City v2 advances satellite-to-city generation from rendering-oriented 3D proxies to explicit textured mesh assets, supported by, to the best of our knowledge, the first documented satellite-mesh paired dataset collected from matched geographic crops for this asset-level task. Project page: https://ai4city-hkust.github.io/Sat2City-v2/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。