通过实例轮廓对齐提升城市空拍定位精度,尤其在密集建筑区表现优异。
LoD-Loc v3: Generalized Aerial Localization in Dense Cities using Instance Silhouette Alignment
- 用新合成数据生成管道构建10万张带实例标注的航拍图
- 实例轮廓对齐使密集区域定位误差大幅降低,零样本泛化能力更强
- 适合需要跨场景高精度空拍定位的研究者与应用开发者
我们提出LoD-Loc v3,一种新型通用空中视觉定位方法,适用于密集城市环境。尽管先前的LoD-Loc v2通过低细节城市模型与语义建筑轮廓对齐实现定位,但仍存在跨场景泛化差、密集建筑区频繁失效的问题。本方法通过两项关键创新解决:首先,开发新的合成数据生成流程,构建InsLoD-Loc——迄今最大的航拍图像实例分割数据集,包含10万张图像及精确的建筑实例标注,使训练模型具备出色的零样本泛化能力;其次,将定位范式从语义轮廓对齐重构为实例轮廓对齐,显著降低密集场景中的位姿估计歧义。大量实验表明,LoD-Loc v3在跨场景与密集城市场景中均显著优于现有最先进基准,性能提升明显。项目地址:https://nudt-sawlab.github.io/LoD-Locv3/
原文摘要 · Abstract (English)
We present LoD-Loc v3, a novel method for generalized aerial visual localization in dense urban environments. While prior work LoD-Loc v2 achieves localization through semantic building silhouette alignment with low-detail city models, it suffers from two key limitations: poor cross-scene generalization and frequent failure in dense building scenes. Our method addresses these challenges through two key innovations. First, we develop a new synthetic data generation pipeline that produces InsLoD-Loc - the largest instance segmentation dataset for aerial imagery to date, comprising 100k images with precise instance building annotations. This enables trained models to exhibit remarkable zero-shot generalization capability. Second, we reformulate the localization paradigm by shifting from semantic to instance silhouette alignment, which significantly reduces pose estimation ambiguity in dense scenes. Extensive experiments demonstrate that LoD-Loc v3 outperforms existing state-of-the-art (SOTA) baselines, achieving superior performance in both cross-scene and dense urban scenarios with a large margin. The project is available at https://nudt-sawlab.github.io/LoD-Locv3/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。