arXiv:2507.11653cs.CVcs.RO2025-07被引 1

无需训练即可在不同视角和季节下实现精准定位

VISTA: Monocular Segmentation-Based Mapping for Appearance and View-Invariant Global Localization

  • 基于物体分割与跟踪构建鲁棒的前端,通过子地图匹配对齐参考帧
  • 在多季节、斜视航拍数据上召回率提升69%,内存仅占基线0.6%
  • 适用于资源受限设备,适合真实场景下的无人导航系统

全局定位对自主导航至关重要,尤其在代理需在不同会话或由其他代理生成的地图中定位时,因缺乏参考帧间的先验关联信息而面临挑战。在非结构化环境中,视角变化、季节更替、空间混淆和遮挡等因素导致传统场景识别方法失效。为此,我们提出VISTA(视不变分割式帧对齐追踪),一种新型开集单目全局定位框架,包含:1)基于物体的分割与跟踪前端;2)利用环境地图间几何一致性进行子地图对应搜索,以对齐车辆参考帧。VISTA可在不同视角和季节变化下保持一致定位性能,无需领域特定训练或微调。我们在多季节及斜视航拍数据集上评估,召回率相比基线最高提升69%。此外,其构建的物体级地图仅占最节省内存基线的0.6%,支持在资源受限平台实现实时运行。

原文摘要 · Abstract (English)

Global localization is critical for autonomous navigation, particularly in scenarios where an agent must localize within a map generated in a different session or by another agent, as agents often have no prior knowledge about the correlation between reference frames. However, this task remains challenging in unstructured environments due to appearance changes induced by viewpoint variation, seasonal changes, spatial aliasing, and occlusions -- known failure modes for traditional place recognition methods. To address these challenges, we propose VISTA (View-Invariant Segmentation-Based Tracking for Frame Alignment), a novel open-set, monocular global localization framework that combines: 1) a front-end, object-based, segmentation and tracking pipeline, followed by 2) a submap correspondence search, which exploits geometric consistencies between environment maps to align vehicle reference frames. VISTA enables consistent localization across diverse camera viewpoints and seasonal changes, without requiring any domain-specific training or finetuning. We evaluate VISTA on seasonal and oblique-angle aerial datasets, achieving up to a 69% improvement in recall over baseline methods. Furthermore, we maintain a compact object-based map that is only 0.6% the size of the most memory-conservative baseline, making our approach capable of real-time implementation on resource-constrained platforms.

全局定位单目视觉物体分割无人机导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。