用卫星图辅助地面照片,快速生成高质量3D场景视图。
Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Images

- 融合卫星图与地面照片,统一坐标系预测高精度高斯点阵。
- 相比仅用地面图像,新视角合成效果显著提升,覆盖更广。
- 适合大规模户外场景重建,尤其适用于数据采集困难的区域。
我们提出Cross-View Splatter,一种前馈式方法,可对地面与卫星拍摄的室外场景同时生成像素对齐的高斯点阵。真实重建需要良好的相机覆盖,但地面影像获取耗时且难以大规模采集。幸运的是,卫星影像可通过公开API轻松获取全局几何先验。该方法将正射校准的卫星视图与地理标记的地面照片融合,在统一的三维坐标系中预测高斯点阵。通过对齐地面与俯视特征表示,模型提升了场景覆盖率和新视角合成质量,优于仅使用地面图像的方法。我们在精选的地理标记数据集及从开放地图服务中挖掘的配对卫星-地形数据上进行训练。在新的基于地理标记影像的新视角合成基准上评估,可与现有最先进方法对比。代码与数据准备流程将公开于 https://nianticspatial.github.io/cross-view-splatter/。
原文摘要 · Abstract (English)
We present Cross-View Splatter, a feed-forward method that predicts pixel-aligned Gaussian splats for outdoor scenes captured at ground level AND by satellite. Faithful reconstructions require good camera coverage, but ground imagery is time-consuming and hard to capture at scale for large outdoor scenes. Fortunately, satellite imagery can provide a global geometric prior that is easy to access via public APIs. Cross-View Splatter fuses orthorectified satellite views with GPS-tagged ground photos to predict Gaussian splats in a unified 3D coordinate frame. By aligning ground and bird's-eye feature representations, our model improves scene coverage and novel-view synthesis, compared to ground imagery alone. We train on curated georeferenced datasets and paired satellite-terrain data, mined from open mapping services. We evaluate our method on a new benchmark for novel-view synthesis with georeferenced imagery allowing comparison to prior state-of-the-art methods. Our code and data preparation will be available at https://nianticspatial.github.io/cross-view-splatter/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。