arXiv:2604.02003cs.CV2026-04

用渐进式扩散模型,从航拍图生成逼真地面视图和3D场景。

ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction

论文配图:ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction
图 1 · 摘自论文原文
  • 分阶段用扩散模型引导高斯点云,逐步逼近地面视角。
  • 在合成与真实数据上均显著提升视觉质量与几何一致性。
  • 无需地面实测数据,适合复杂场景重建任务。

仅基于航拍图像生成地面视图和一致的3D场景模型极具挑战,原因在于视角变化剧烈、中间观测缺失以及尺度差异大。现有方法或在渲染后进行修正,常导致几何不一致;或依赖多高度地面真值,而此类数据极少可得。高斯点阵与基于扩散的优化虽在小范围变化下表现良好,但在宽幅航拍到地面视角跳跃时失效。为此,本文提出ProDiG(Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction),一种渐进式扩散引导框架,逐步将航拍3D表示转换为地面级保真度。ProDiG合成中间高度视图,并在每阶段利用几何感知因果注意力模块,将对极结构注入参考视图的扩散过程。距离自适应高斯模块根据相机距离动态调整高斯点尺度与不透明度,确保在大视角跨度下稳定重建。二者协同实现无需额外地面真值的渐进式、几何约束的优化。在合成与真实数据集上的大量实验表明,ProDiG生成的地面渲染图像视觉真实感强,3D几何连贯,显著优于现有方法,在视觉质量、几何一致性及极端视角变化鲁棒性方面均有提升。

原文摘要 · Abstract (English)

Generating ground-level views and coherent 3D site models from aerial-only imagery is challenging due to extreme viewpoint changes, missing intermediate observations, and large scale variations. Existing methods either refine renderings post-hoc, often producing geometrically inconsistent results, or rely on multi-altitude ground-truth, which is rarely available. Gaussian Splatting and diffusion-based refinements improve fidelity under small variations but fail under wide aerial-toground gaps. To address these limitations, we introduce ProDiG (Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction), a diffusionguided framework that progressively transforms aerial 3D representations toward ground-level fidelity. ProDiG synthesizes intermediate-altitude views and refines the Gaussian representation at each stage using a geometry-aware causal attention module that injects epipolar structure into reference-view diffusion. A distance-adaptive Gaussian module dynamically adjusts Gaussian scale and opacity based on camera distance, ensuring stable reconstruction across large viewpoint gaps. Together, these components enable progressive, geometrically grounded refinement without requiring additional ground-truth viewpoints. Extensive experiments on synthetic and real-world datasets demonstrate that ProDiG produces visually realistic ground-level renderings and coherent 3D geometry, significantly outperforming existing approaches in terms of visual quality, geometric consistency, and robustness to extreme viewpoint changes. Project Page: https://sirsh07.github.io/research/prodig

三维重建扩散模型航拍转地面高斯点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。