arXiv:2412.03315cs.CV2024-12被引 8

用几何引导扩散模型,实现卫星与街景间一对多图像生成

Geometry-guided Cross-view Diffusion for One-to-many Cross-view Image Synthesis

  • 引入几何对齐条件,解决视角间位置模糊问题
  • 在三个数据集上生成图像质量、保真度和多样性均更优
  • 适合需要多视角真实图像生成的研究者或应用开发

本文提出一种新型跨视角图像合成方法,旨在从卫星图像生成对应的地面视图,或反之。我们称其为卫星到地面(Sat2Grd)与地面到卫星(Grd2Sat)合成任务。不同于以往仅支持一对一生成的方法,本方法认识到该问题固有的“一对多”特性——由于光照、天气及遮挡差异导致的输出不确定性。为此,我们利用扩散模型中的随机高斯噪声来表征目标视图数据中学习到的多种可能性,并设计了几何引导的跨视角条件(GCC)策略,建立卫星与街景特征间的显式几何对应关系,从而缓解由相机姿态差异引入的几何模糊性。在三个基准跨视角数据集上的大量定量与定性分析表明,所提方法显著优于基线模型,包括当前最先进的跨视角图像合成方法,在图像质量、保真度和多样性方面表现更佳。

原文摘要 · Abstract (English)

This paper presents a novel approach for cross-view synthesis aimed at generating plausible ground-level images from corresponding satellite imagery or vice versa. We refer to these tasks as satellite-to-ground (Sat2Grd) and ground-to-satellite (Grd2Sat) synthesis, respectively. Unlike previous works that typically focus on one-to-one generation, producing a single output image from a single input image, our approach acknowledges the inherent one-to-many nature of the problem. This recognition stems from the challenges posed by differences in illumination, weather conditions, and occlusions between the two views. To effectively model this uncertainty, we leverage recent advancements in diffusion models. Specifically, we exploit random Gaussian noise to represent the diverse possibilities learnt from the target view data. We introduce a Geometry-guided Cross-view Condition (GCC) strategy to establish explicit geometric correspondences between satellite and street-view features. This enables us to resolve the geometry ambiguity introduced by camera pose between image pairs, boosting the performance of cross-view image synthesis. Through extensive quantitative and qualitative analyses on three benchmark cross-view datasets, we demonstrate the superiority of our proposed geometry-guided cross-view condition over baseline methods, including recent state-of-the-art approaches in cross-view image synthesis. Our method generates images of higher quality, fidelity, and diversity than other state-of-the-art approaches.

图像生成扩散模型跨视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。