arXiv:2511.08258cs.CV2025-11

用高空图生成地面视角图,不依赖深度或3D模型

Top2Ground: A Height-Aware Dual Conditioning Diffusion Model for Robust Aerial-to-Ground View Generation

  • 结合空间特征与语义信息进行双重条件扩散
  • 在三个数据集上平均SSIM提升7.3%
  • 适合需要跨视角图像生成的导航与遥感应用

从高空视图生成地面视角图像是一项挑战性任务,受限于视角差异、遮挡和视野范围。我们提出Top2Ground,一种基于扩散模型的新方法,直接从高空RGB图像生成逼真的地面视角图像,无需依赖深度图或3D体素等中间表示。具体而言,通过联合使用VAE编码的空间特征(来自高空RGB图像和估计的高度图)以及CLIP生成的语义嵌入,对去噪过程进行条件控制。该设计使生成结果既符合场景的三维结构几何约束,又保持内容语义一致性。我们在三个多样化数据集(CVUSA、CVACT、Auto Arborist)上评估该方法,结果显示在三个基准数据集上平均SSIM提升7.3%,表明该方法能有效应对宽窄不同视野,展现出强大的泛化能力。

原文摘要 · Abstract (English)

Generating ground-level images from aerial views is a challenging task due to extreme viewpoint disparity, occlusions, and a limited field of view. We introduce Top2Ground, a novel diffusion-based method that directly generates photorealistic ground-view images from aerial input images without relying on intermediate representations such as depth maps or 3D voxels. Specifically, we condition the denoising process on a joint representation of VAE-encoded spatial features (derived from aerial RGB images and an estimated height map) and CLIP-based semantic embeddings. This design ensures the generation is both geometrically constrained by the scene's 3D structure and semantically consistent with its content. We evaluate Top2Ground on three diverse datasets: CVUSA, CVACT, and the Auto Arborist. Our approach shows 7.3% average improvement in SSIM across three benchmark datasets, showing Top2Ground can robustly handle both wide and narrow fields of view, highlighting its strong generalization capabilities.

图像生成扩散模型跨视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。