arXiv:2605.14984cs.CVcs.AI2026-05被引 5

从一张卫星图生成高精度街景3D场景,几何更准、视觉更真。

Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image

论文配图:Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image
图 1 · 摘自论文原文
  • 采用几何优先策略,结合视角训练和新约束提升重建精度。
  • 几何误差降低至5.20米(原6.76米),图像真实感FID降至19。
  • 适合需要高质量3D城市资产的科研与工程应用。

从单张卫星图像生成街景级3D场景是一项关键但极具挑战的任务。现有方法存在明显权衡:几何-色彩化模型虽几何精度高,但仅聚焦建筑且语义单一;代理模型通过端到端图像到3D框架联合学习几何与纹理,内容丰富但几何粗糙不稳。我们归因于卫星到街景视角差距大、监督稀疏且不一致。为此提出Sat3DGen,采用几何优先方法,通过引入新型几何约束与视角视图训练策略,有效缓解主要几何误差来源。该方法显著提升3D准确度与逼真度。我们构建新基准,将VIGOR-OOD测试集与高分辨率DSM数据配对。在该基准上,几何均方根误差由6.76米降至5.20米。更重要的是,几何提升带动视觉真实感飞跃,即使未使用额外图像质量模块,弗雷切特起始距离(FID)从约40降至19,优于领先方法Sat2Density++。我们展示了高质量3D资产在语义地图转3D、多相机视频生成、大规模网格化及无监督单图DSM估计等下游任务中的多样性应用。代码已开源于https://github.com/qianmingduowan/Sat3DGen。

原文摘要 · Abstract (English)

Generating a street-level 3D scene from a single satellite image is a crucial yet challenging task. Current methods present a stark trade-off: geometry-colorization models achieve high geometric fidelity but are typically building-focused and lack semantic diversity. In contrast, proxy-based models use feed-forward image-to-3D frameworks to generate holistic scenes by jointly learning geometry and texture, a process that yields rich content but coarse and unstable geometry. We attribute these geometric failures to the extreme viewpoint gap and sparse, inconsistent supervision inherent in satellite-to-street data. We introduce Sat3DGen to address these fundamental challenges, which embodies a geometry-first methodology. This methodology enhances the feed-forward paradigm by integrating novel geometric constraints with a perspective-view training strategy, explicitly countering the primary sources of geometric error. This geometry-centric strategy yields a dramatic leap in both 3D accuracy and photorealism. For validation, we first constructed a new benchmark by pairing the VIGOR-OOD test set with high-resolution DSM data. On this benchmark, our method improves geometric RMSE from 6.76m to 5.20m. Crucially, this geometric leap also boosts photorealism, reducing the Fréchet Inception Distance (FID) from $\sim$40 to 19 against the leading method, Sat2Density++, despite using no extra tailored image-quality modules. We demonstrate the versatility of our high-quality 3D assets through diverse downstream applications, including semantic-map-to-3D synthesis, multi-camera video generation, large-scale meshing, and unsupervised single-image Digital Surface Model (DSM) estimation. The code has been released on https://github.com/qianmingduowan/Sat3DGen.

3D生成卫星图像几何优化城市建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。