构建百万级航拍数据集并用旋转不变网络提升无人机定位精度
Visual Place Recognition for Large-Scale UAV Applications
- 构建包含100万张图像的LASED航拍数据集,覆盖埃塞俄比亚十年间17万处地点
- 使用可旋转卷积网络,在航拍图像中实现12%的召回率提升
- 适合研究大规模无人机视觉定位与旋转鲁棒性模型的研究者
视觉场景识别(vPR)在无人机导航中至关重要,可实现跨多样化环境的稳定定位。尽管已有显著进展,航空vPR仍面临挑战:缺乏大规模高空数据集导致模型泛化能力受限,且航拍图像存在固有的旋转模糊性。为此,我们提出LASED——一个包含约一百万张图像的大规模航拍数据集,系统采样自十年间爱沙尼亚17万处独特位置,具备丰富的地理与时间多样性。其结构化设计确保了清晰的场景分离,显著提升航空场景下的模型训练效果。此外,我们引入可旋转卷积神经网络(steerable CNNs),利用其固有的旋转等变性,生成对方向不敏感的特征表示,有效应对航拍图像中的旋转模糊问题。大量基准测试表明,基于LASED训练的模型召回率显著高于在小规模、低多样性数据集上训练的模型,凸显广泛地理覆盖与时间多样性的优势。同时,可旋转CNN相比传统卷积架构持续表现更优,平均召回率提升12%。结合结构化大规模数据集与旋转等变神经网络,该方法显著增强了航空vPR中模型的鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Visual Place Recognition (vPR) plays a crucial role in Unmanned Aerial Vehicle (UAV) navigation, enabling robust localization across diverse environments. Despite significant advancements, aerial vPR faces unique challenges due to the limited availability of large-scale, high-altitude datasets, which limits model generalization, along with the inherent rotational ambiguity in UAV imagery. To address these challenges, we introduce LASED, a large-scale aerial dataset with approximately one million images, systematically sampled from 170,000 unique locations throughout Estonia over a decade, offering extensive geographic and temporal diversity. Its structured design ensures clear place separation significantly enhancing model training for aerial scenarios. Furthermore, we propose the integration of steerable Convolutional Neural Networks (CNNs) to explicitly handle rotational variance, leveraging their inherent rotational equivariance to produce robust, orientation-invariant feature representations. Our extensive benchmarking demonstrates that models trained on LASED achieve significantly higher recall compared to those trained on smaller, less diverse datasets, highlighting the benefits of extensive geographic coverage and temporal diversity. Moreover, steerable CNNs effectively address rotational ambiguity inherent in aerial imagery, consistently outperforming conventional convolutional architectures, achieving on average 12\% recall improvement over the best-performing non-steerable network. By combining structured, large-scale datasets with rotation-equivariant neural networks, our approach significantly enhances model robustness and generalization for aerial vPR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。