用3D几何重建统一无人机跨视角定位,提升城市环境定位精度。
Unifying UAV Cross-View Geo-Localization via 3D Geometric Perception
- 通过视觉几何变压器重建3D场景,生成对齐卫星图的鸟瞰视图
- 实现端到端米级定位,在复杂城市区域表现更优
- 适合需要高精度定位的无人机导航与自动驾驶系统
在无全球导航卫星系统(GNSS)环境下,无人机从倾斜影像到正射卫星图的严重几何差异仍使跨视角地理定位面临挑战。现有方法多采用先检索后估计的分离式流程,将透视畸变视为外观噪声而非显式几何变换。本文提出一种几何感知的无人机定位框架,通过视觉几何基础变换器(VGGT)从多视角无人机图像序列重建局部3D场景,并渲染虚拟鸟瞰视图(BEV),对齐卫星影像。该BEV作为几何中介,支持鲁棒跨视角检索并提供空间先验以实现3自由度(3-DoF)姿态回归。为高效处理多个位置假设,引入卫星注意力模块,隔离卫星候选与重建场景间的交互,保持线性计算复杂度。同时发布经过坐标精标与空间重叠分析重构的University-1652数据集,支持端到端定位评估。在优化后的University-1652与SUES-200基准上,本方法显著优于现有基线,实现稳健的米级定位精度,并在复杂城市环境中展现更强泛化能力。
原文摘要 · Abstract (English)
Cross-view geo-localization for Unmanned Aerial Vehicles (UAVs) operating in GNSS-denied environments remains challenging due to the severe geometric discrepancy between oblique UAV imagery and orthogonal satellite maps. Most existing methods address this problem through a decoupled pipeline of place retrieval and pose estimation, implicitly treating perspective distortion as appearance noise rather than an explicit geometric transformation. In this work, we propose a geometry-aware UAV geo-localization framework that explicitly models the 3D scene geometry to unify coarse place recognition and fine-grained pose estimation within a single inference pipeline. Our approach reconstructs a local 3D scene from multi-view UAV image sequences using a Visual Geometry Grounded Transformer (VGGT), and renders a virtual Bird's-Eye View (BEV) representation that orthorectifies the UAV perspective to align with satellite imagery. This BEV serves as a geometric intermediary that enables robust cross-view retrieval and provides spatial priors for accurate 3 Degrees of Freedom (3-DoF) pose regression. To efficiently handle multiple location hypotheses, we introduce a Satellite-wise Attention Block that isolates the interaction between each satellite candidate and the reconstructed UAV scene, preventing inter-candidate interference while maintaining linear computational complexity. In addition, we release a recalibrated version of the University-1652 dataset with precise coordinate annotations and spatial overlap analysis, enabling rigorous evaluation of end-to-end localization accuracy. Extensive experiments on the refined University-1652 benchmark and SUES-200 demonstrate that our method significantly outperforms state-of-the-art baselines, achieving robust meter-level localization accuracy and improved generalization in complex urban environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。