arXiv:2608.21926cs.CV2026-08

用视觉几何特征提升无人机最后十米导航的定位精度

AirAlign: Geometry-Aware Relative Pose Alignment for UAV Last-Meter Navigation

论文配图:AirAlign: Geometry-Aware Relative Pose Alignment for UAV Last-Meter Navigation
图 1 · 摘自论文原文
  • 基于预训练模型提取图像几何特征,实现无深度信息下的精准位姿对齐
  • 在PairUAV数据集上达到0.81°的旋转误差和1.23米的平移误差
  • 适合需要高精度低空导航的无人机应用,如精准投送或设备操作

现代低空环境中无人机导航在最终接近阶段需要更高精度的位姿对齐,以完成目标信息获取或操作,因此“最后十米”导航愈发关键。然而,视角和外观剧烈变化使该任务极具挑战。为此,我们提出AirAlign,一种仅依赖RGB图像对的无人机相对位姿对齐框架。AirAlign采用预训练的视觉几何重建模型作为主干网络,从源-目标图像对中提取几何感知特征。为更高效利用有限训练数据,我们将训练集划分为多个场景互不重叠的子集,进行未见场景交叉验证与模型选择。推理时,选取的多个模型预测结果取平均,形成整体集成输出。在ACMMM 2026无人机多媒体研讨会的PairUAV挑战赛上的实验表明,该方法具有优异的性能与鲁棒性,全面的消融研究也验证了各组件的有效贡献。

原文摘要 · Abstract (English)

Unmanned aerial vehicle (UAV) navigation in modern low-altitude environments requires more accurate pose alignment in the final approach stage for target information acquisition or manipulation, making "last-meter" navigation increasingly important. However, severe viewpoint and appearance variations make this task challenging. To tackle this problem, we propose AirAlign, a framework for RGB-only image-pair relative pose alignment for UAVs. AirAlign uses a pretrained visual geometry reconstruction model as the backbone to extract geometry-aware features from source-target image pairs. In addition, to better utilize the limited training data, we split the training set into multiple scene-disjoint folds for unseen cross-validation and model selection. During inference, the predictions of the selected models are averaged to form the ensemble output of the overall framework. Experiments on the PairUAV challenge at the ACMMM 2026 Workshop on UAVs in Multimedia demonstrate the effectiveness and robustness of our method, while comprehensive ablation studies validate the contribution of each component.

无人机导航位姿估计视觉几何图像对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。