用语义3D城市模型实现无人机厘米级定位,无需高精度卫星信号。
SemCityLoc: Aerial 6DoF Localization Using Semantic 3D City Models

- 通过语义表面与单目深度对齐,提升重复和遮挡环境下的定位区分度。
- 在城市峡谷中定位误差从9.89米降至2.62米,召回率提升36%。
- 适合无人机、自动驾驶等需要高精度低资源定位的场景。
航空6自由度定位通常依赖高精度GNSS信号或辐射丰富的3D重建,限制了可扩展性和机载部署。本文提出SemCityLoc,一种基于语义-几何对齐的系统,将航拍姿态估计重构为由基础模型生成的视觉先验与标准化LoD合规3D城市模型之间的结构化表面配准问题。不同于匹配稀疏轮廓或密集纹理,本方法对齐语义表面与单目深度,结合轻量级语义3D建筑模型,在重复和遮挡的城市环境中显著提升姿态可辨识性。为支持精确评估,我们引入SemCityLockeD,首个真实世界基准,融合厘米级无人机姿态与标准化LoD1–LoD3语义城市模型及挑战性的低空影像。实验表明,相比现有地图基方法,定位召回率最高提升36%,城市峡谷中的平均位置误差由9.89米降至2.62米。结果表明,语义结构化的几何信息足以提供高精度航空定位所需的充足且可扩展约束,无需辐射场景重建。代码与数据见https://albertchen98.github.io/SemCityLoc。
原文摘要 · Abstract (English)
Aerial 6DoF localization typically relies on precise GNSS signals or radiometrically rich 3D reconstructions, limiting scalability and on-board deployment. We propose SemCityLoc, a semantic-geometric alignment system that reframes aerial pose estimation as structured surface registration between foundation-model-derived visual priors and standardized LoD-compliant 3D city models. Instead of matching sparse contours or dense texture, our method aligns semantic surfaces and monocular depth with lightweight semantic 3D building models, increasing pose discriminability in repetitive and occluded urban environments. To enable accurate evaluation, we introduce SemCityLockeD, the first real-world benchmark combining centimeter-accurate UAV poses with standardized LoD1--LoD3 semantic city models and challenging low-altitude imagery. Experiments demonstrate substantial improvements over existing map-based approaches, improving recall by up to 36% and reducing mean positional error from 9.89m to 2.62m in challenging urban canyons. Our results indicate that semantically structured geometry provides sufficient and scalable constraints for high-precision aerial localization without radiometric scene reconstructions. The code and data are available at https://albertchen98.github.io/SemCityLoc.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。