arXiv:2507.00659cs.CV2025-07ICCV被引 6

让无人机在低精度城市模型上实现精准定位,首次突破低细节模型限制。

LoD-Loc v2: Aerial Visual Localization over Low Level-of-Detail City Models using Explicit Silhouette Alignment

  • 通过轮廓对齐的粗细结合策略,提升空中定位精度。
  • 在10.7公里范围内实现高精度定位,支持大初始误差收敛。
  • 开源两个LoD1数据集,推动全球城市无人机定位研究。

我们提出一种新型方法,用于在低细节层级(LoD1)城市模型上实现航拍视觉定位。先前基于线框对齐的方法LoD-Loc虽有良好表现,但主要依赖高细节层级(LoD3或LoD2)模型,而多数现有及计划建设的全国性城市模型均为低细节层级(LoD1)。因此,在低细节模型上实现定位可释放无人机在全球城市环境中的潜力。为此,我们提出LoD-Loc v2,采用显式轮廓对齐的粗到精策略,在空中实现低细节模型上的精准定位。给定查询图像,首先使用建筑分割网络提取建筑轮廓;在粗略姿态选择阶段,通过均匀采样先验姿态周围的姿态假设构建姿态代价体,每个代价衡量投影轮廓与预测轮廓的对齐程度,选取最大值对应姿态作为粗略姿态;在精细姿态估计阶段,采用融合多束跟踪的粒子滤波方法高效探索假设空间,获得最终姿态估计。为促进该领域研究,我们发布了两个覆盖10.7公里范围的LoD1城市模型数据集,包含真实RGB查询图像与真值姿态标注。实验表明,LoD-Loc v2不仅在高细节模型上提升精度,更首次实现低细节模型上的定位,并显著优于当前最优基线,甚至超越纹理模型方法,且扩大了收敛区域以适应更大初始误差。

原文摘要 · Abstract (English)

We propose a novel method for aerial visual localization over low Level-of-Detail (LoD) city models. Previous wireframe-alignment-based method LoD-Loc has shown promising localization results leveraging LoD models. However, LoD-Loc mainly relies on high-LoD (LoD3 or LoD2) city models, but the majority of available models and those many countries plan to construct nationwide are low-LoD (LoD1). Consequently, enabling localization on low-LoD city models could unlock drones' potential for global urban localization. To address these issues, we introduce LoD-Loc v2, which employs a coarse-to-fine strategy using explicit silhouette alignment to achieve accurate localization over low-LoD city models in the air. Specifically, given a query image, LoD-Loc v2 first applies a building segmentation network to shape building silhouettes. Then, in the coarse pose selection stage, we construct a pose cost volume by uniformly sampling pose hypotheses around a prior pose to represent the pose probability distribution. Each cost of the volume measures the degree of alignment between the projected and predicted silhouettes. We select the pose with maximum value as the coarse pose. In the fine pose estimation stage, a particle filtering method incorporating a multi-beam tracking approach is used to efficiently explore the hypothesis space and obtain the final pose estimation. To further facilitate research in this field, we release two datasets with LoD1 city models covering 10.7 km , along with real RGB queries and ground-truth pose annotations. Experimental results show that LoD-Loc v2 improves estimation accuracy with high-LoD models and enables localization with low-LoD models for the first time. Moreover, it outperforms state-of-the-art baselines by large margins, even surpassing texture-model-based methods, and broadens the convergence basin to accommodate larger prior errors.

视觉定位无人机城市建模轮廓对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。