用多张地面照片重建3D场景,精准定位到卫星图上。
Wrivinder: Towards Spatial Intelligence for Geo-locating Ground Images onto Satellite Imagery
- 通过多视角照片重建3D场景,结合几何与语义信息对齐卫星图。
- 零样本下在密集和大范围场景中实现30米内定位精度。
- 适合做地图、导航或无GPS环境下的定位研究者使用。
将地面影像与地理注册的卫星地图对齐对于地图绘制、导航和态势感知至关重要,但在视角差异大或GPS不可靠时仍具挑战。我们提出Wrivinder,一种零样本、基于几何的框架,通过聚合多张地面照片重建一致的3D场景,并与高空卫星影像对齐。Wrivinder融合SfM重建、3D高斯泼溅、语义定位及单目深度度量线索,生成可直接匹配卫星上下文的俯视渲染图,实现度量准确的相机地理定位。为系统评估该任务(目前缺乏合适基准),我们还发布了MC-Sat数据集,包含跨多样化户外环境的多视角地面影像与地理注册卫星瓦片。两者共同构建了首个无需成对监督的几何驱动跨视图对齐基准。零样本实验中,Wrivinder在密集与大范围场景中均达到30米以内的地理定位精度,凸显几何聚合在鲁棒地面-卫星定位中的潜力。
原文摘要 · Abstract (English)
Aligning ground-level imagery with geo-registered satellite maps is crucial for mapping, navigation, and situational awareness, yet remains challenging under large viewpoint gaps or when GPS is unreliable. We introduce Wrivinder, a zero-shot, geometry-driven framework that aggregates multiple ground photographs to reconstruct a consistent 3D scene and align it with overhead satellite imagery. Wrivinder combines SfM reconstruction, 3D Gaussian Splatting, semantic grounding, and monocular depth--based metric cues to produce a stable zenith-view rendering that can be directly matched to satellite context for metrically accurate camera geo-localization. To support systematic evaluation of this task, which lacks suitable benchmarks, we also release MC-Sat, a curated dataset linking multi-view ground imagery with geo-registered satellite tiles across diverse outdoor environments. Together, Wrivinder and MC-Sat provide a first comprehensive baseline and testbed for studying geometry-centered cross-view alignment without paired supervision. In zero-shot experiments, Wrivinder achieves sub-30\,m geolocation accuracy across both dense and large-area scenes, highlighting the promise of geometry-based aggregation for robust ground-to-satellite localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。