arXiv:2509.24369cs.CVcs.AI2025-09

融合扩散模型与GAN,从卫星图生成逼真街景。

From Satellite to Street: A Hybrid Framework Integrating Stable Diffusion and PanoGAN for Consistent Cross-View Synthesis

  • 用双分支架构结合Stable Diffusion与条件GAN生成街景。
  • 在CVUSA数据集上优于纯扩散模型,保持几何一致性。
  • 适合城市分析、自动驾驶等需跨视角图像的场景。

街景图像已成为地理空间数据采集和城市分析的重要来源,支持决策制定。然而,从卫星图像生成街景面临外观和视角差异的巨大挑战。本文提出一种混合框架,结合基于扩散的模型与条件生成对抗网络,实现从卫星图像生成地理一致的街景图像。该方法采用多阶段训练策略,以Stable Diffusion为核心组件,构建双分支架构,并引入条件生成对抗网络(GAN),生成地理一致的全景街景。通过融合两种模型的优势,提升生成图像的几何一致性和视觉质量。在具有挑战性的跨视角美国数据集(Cross-View USA, CVUSA)上进行评估,实验结果表明,该混合方法在多个评价指标上优于纯扩散模型,性能与最先进的基于GAN的方法相当。框架成功生成真实且几何一致的街景图像,保留了道路标记、次级道路及云层等细粒度局部细节。

原文摘要 · Abstract (English)

Street view imagery has become an essential source for geospatial data collection and urban analytics, enabling the extraction of valuable insights that support informed decision-making. However, synthesizing street-view images from corresponding satellite imagery presents significant challenges due to substantial differences in appearance and viewing perspective between these two domains. This paper presents a hybrid framework that integrates diffusion-based models and conditional generative adversarial networks to generate geographically consistent street-view images from satellite imagery. Our approach uses a multi-stage training strategy that incorporates Stable Diffusion as the core component within a dual-branch architecture. To enhance the framework's capabilities, we integrate a conditional Generative Adversarial Network (GAN) that enables the generation of geographically consistent panoramic street views. Furthermore, we implement a fusion strategy that leverages the strengths of both models to create robust representations, thereby improving the geometric consistency and visual quality of the generated street-view images. The proposed framework is evaluated on the challenging Cross-View USA (CVUSA) dataset, a standard benchmark for cross-view image synthesis. Experimental results demonstrate that our hybrid approach outperforms diffusion-only methods across multiple evaluation metrics and achieves competitive performance compared to state-of-the-art GAN-based methods. The framework successfully generates realistic and geometrically consistent street-view images while preserving fine-grained local details, including street markings, secondary roads, and atmospheric elements such as clouds.

图像生成跨视角扩散模型城市分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。