用语义与结构约束实现无人机视觉定位的分层图像匹配
Hierarchical Image Matching for UAV Absolute Visual Localization via Semantic and Structural Constraints
- 分两阶段匹配:先语义约束粗匹配,再轻量细粒度精匹配
- 在公开数据集和新自建CS-UAV数据集上定位误差更小、鲁棒性更强
- 适合无GPS环境下依赖视觉定位的无人机系统
绝对定位旨在确定智能体相对于全局参考系的位置,对无人机在多种场景下的应用至关重要,但在全球导航卫星系统(GNSS)信号不可用时面临挑战。基于视觉的绝对定位方法通过将当前无人机视角与参考卫星地图进行匹配来估计位置,已成为无GNSS环境下的主流方案。然而,现有方法多依赖传统低层次图像匹配,在跨源差异和时间变化带来的显著差异面前表现不佳。为此,本文提出一种面向无人机绝对视觉定位的分层跨源图像匹配方法,结合语义感知与结构约束的粗匹配模块和轻量级细粒度匹配模块。在粗匹配阶段,利用视觉基础模型提取的语义特征,在语义与结构约束下建立区域级对应关系;随后在细粒度匹配阶段提取精细特征并建立像素级对应。基于此,构建了一套不依赖相对定位技术的无人机绝对视觉定位流程,主要通过图像检索模块前置实现。在公开基准数据集和新提出的CS-UAV数据集上的实验验证了该方法在多种复杂条件下的优越精度与鲁棒性,证明其有效性。
原文摘要 · Abstract (English)
Absolute localization, aiming to determine an agent's location with respect to a global reference, is crucial for unmanned aerial vehicles (UAVs) in various applications, but it becomes challenging when global navigation satellite system (GNSS) signals are unavailable. Vision-based absolute localization methods, which locate the current view of the UAV in a reference satellite map to estimate its position, have become popular in GNSS-denied scenarios. However, existing methods mostly rely on traditional and low-level image matching, suffering from difficulties due to significant differences introduced by cross-source discrepancies and temporal variations. To overcome these limitations, in this paper, we introduce a hierarchical cross-source image matching method designed for UAV absolute localization, which integrates a semantic-aware and structure-constrained coarse matching module with a lightweight fine-grained matching module. Specifically, in the coarse matching module, semantic features derived from a vision foundation model first establish region-level correspondences under semantic and structural constraints. Then, the fine-grained matching module is applied to extract fine features and establish pixel-level correspondences. Building upon this, a UAV absolute visual localization pipeline is constructed without any reliance on relative localization techniques, mainly by employing an image retrieval module before the proposed hierarchical image matching modules. Experimental evaluations on public benchmark datasets and a newly introduced CS-UAV dataset demonstrate superior accuracy and robustness of the proposed method under various challenging conditions, confirming its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。