arXiv:2508.01936cs.CVcs.RO2025-08被引 2

跨视角深度框架提升多高度稀疏定位精度与鲁棒性

CVD-SfM: A Cross-View Deep Front-end Structure-from-Motion System for Sparse Localization in Multi-Altitude Scenes

  • 融合跨视图变换器与深度特征的统一结构光流系统
  • 在新构建数据集上显著优于现有方法,定位更准更稳
  • 适合无人机导航、搜救与巡检等真实场景应用

我们提出一种新型多高度相机位姿估计算法,针对仅依赖稀疏图像输入时在不同高度下实现鲁棒且精确定位的挑战。通过将跨视图变换器、深度特征与运动恢复结构(SfM)整合至统一框架,有效应对多样环境条件与视角变化。为评估方法并推动研究,我们构建了两个专用于多高度位姿估计的新数据集,此类数据集在当前文献中极为罕见。所提框架在这些数据集上经过大量对比实验验证,结果表明其在多高度稀疏位姿估计任务中,精度与鲁棒性均优于现有方案,适用于无人机导航、搜救及自动化巡检等实际机器人应用。

原文摘要 · Abstract (English)

We present a novel multi-altitude camera pose estimation system, addressing the challenges of robust and accurate localization across varied altitudes when only considering sparse image input. The system effectively handles diverse environmental conditions and viewpoint variations by integrating the cross-view transformer, deep features, and structure-from-motion into a unified framework. To benchmark our method and foster further research, we introduce two newly collected datasets specifically tailored for multi-altitude camera pose estimation; datasets of this nature remain rare in the current literature. The proposed framework has been validated through extensive comparative analyses on these datasets, demonstrating that our system achieves superior performance in both accuracy and robustness for multi-altitude sparse pose estimation tasks compared to existing solutions, making it well suited for real-world robotic applications such as aerial navigation, search and rescue, and automated inspection.

位姿估计多高度稀疏定位无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。