arXiv:2504.13157cs.CV2025-04CVPR被引 42

用合成与真实图像混合数据,提升航拍与地面视角的三维重建精度。

AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis

  • 融合城市3D模型合成航拍图与真实地面图像,构建混合训练数据。
  • 在零样本场景下,相机旋转误差小于5度的匹配率从不足5%提升至56%。
  • 适用于航拍-地面视角转换、新视角生成等实际应用,尤其适合大视角变化场景。

本文研究航拍与地面视角图像的几何重建问题。现有基于学习的方法难以应对二者间极端视角差异。我们提出,缺乏高质量、配准精确的航拍-地面数据集是主要原因。为此,我们设计了一种可扩展框架:利用3D城市网格(如Google Earth)生成伪合成航拍图像,结合真实、众包获取的地面图像(如MegaDepth),以提升地表细节真实性,有效弥合真实图像与合成数据之间的领域差距。基于该混合数据集,我们微调多个先进算法,在真实世界零样本任务中取得显著提升。例如,基线方法DUSt3R在航拍-地面配对中仅能定位少于5%的图像对,旋转误差低于5度;而使用本数据微调后,准确率提升至近56%,显著改善了大视角变化下的处理能力。此外,该数据集还提升了下游任务如复杂航拍-地面场景中的新视角合成性能,证明了其在真实应用中的价值。

原文摘要 · Abstract (English)

We explore the task of geometric reconstruction of images captured from a mixture of ground and aerial views. Current state-of-the-art learning-based approaches fail to handle the extreme viewpoint variation between aerial-ground image pairs. Our hypothesis is that the lack of high-quality, co-registered aerial-ground datasets for training is a key reason for this failure. Such data is difficult to assemble precisely because it is difficult to reconstruct in a scalable way. To overcome this challenge, we propose a scalable framework combining pseudo-synthetic renderings from 3D city-wide meshes (e.g., Google Earth) with real, ground-level crowd-sourced images (e.g., MegaDepth). The pseudo-synthetic data simulates a wide range of aerial viewpoints, while the real, crowd-sourced images help improve visual fidelity for ground-level images where mesh-based renderings lack sufficient detail, effectively bridging the domain gap between real images and pseudo-synthetic renderings. Using this hybrid dataset, we fine-tune several state-of-the-art algorithms and achieve significant improvements on real-world, zero-shot aerial-ground tasks. For example, we observe that baseline DUSt3R localizes fewer than 5% of aerial-ground pairs within 5 degrees of camera rotation error, while fine-tuning with our data raises accuracy to nearly 56%, addressing a major failure point in handling large viewpoint changes. Beyond camera estimation and scene reconstruction, our dataset also improves performance on downstream tasks like novel-view synthesis in challenging aerial-ground scenarios, demonstrating the practical value of our approach in real-world applications.

三维重建航拍图像视角合成数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。