从卫星图生成连贯的地面视角图像,解决视角差异与多图不一致问题。
Satellite to GroundScape -- Large-scale Consistent Ground View Generation from Satellite Views
- 用卫星图引导去噪过程,提取场景布局确保生成一致性。
- 引入时序去噪模块,保持多视角图像间运动连续性。
- 构建超10万对数据集,支持大规模地面场景生成。
从卫星图像生成一致的地面视角图像极具挑战,主要源于卫星与地面视角在拍摄角度和分辨率上的巨大差异。以往工作多集中于单视角生成,常导致相邻地面视角之间不连贯。本文提出一种新型跨视角合成方法,通过固定潜在扩散模型,引入两个条件模块:卫星引导去噪模块,用于提取高层场景布局以指导去噪过程;卫星时序去噪模块,用于捕捉相机运动以维持多视角生成结果的一致性。我们还构建了一个包含超过10万对视角的大型卫星-地面数据集,以支持大规模地面场景或视频生成。实验表明,该方法在感知质量和时序一致性指标上均优于现有方法,生成的多视角输出具有高保真度与高度一致性。
原文摘要 · Abstract (English)
Generating consistent ground-view images from satellite imagery is challenging, primarily due to the large discrepancies in viewing angles and resolution between satellite and ground-level domains. Previous efforts mainly concentrated on single-view generation, often resulting in inconsistencies across neighboring ground views. In this work, we propose a novel cross-view synthesis approach designed to overcome these challenges by ensuring consistency across ground-view images generated from satellite views. Our method, based on a fixed latent diffusion model, introduces two conditioning modules: satellite-guided denoising, which extracts high-level scene layout to guide the denoising process, and satellite-temporal denoising, which captures camera motion to maintain consistency across multiple generated views. We further contribute a large-scale satellite-ground dataset containing over 100,000 perspective pairs to facilitate extensive ground scene or video generation. Experimental results demonstrate that our approach outperforms existing methods on perceptual and temporal metrics, achieving high photorealism and consistency in multi-view outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。