arXiv:2506.05250cs.CVcs.RO2025-06被引 3

无需GPS的野外环境定位,靠自监督模型精准识别视角与季节变化下的场景。

Spatiotemporal Contrastive Learning for Cross-View Video Localization in Unstructured Off-road Terrains

  • 通过姿态相关采样和时序负样本挖掘,提升方向辨别能力。
  • 在12.29公里测试中,93%定位误差小于25米,100%小于50米。
  • 适用于真实复杂地形,且不需针对特定环境微调。

在无GPS、非结构化野外环境中实现鲁棒的跨视角3-DoF定位仍具挑战,主要源于重复植被和无序地形导致的感知歧义,以及季节变化对场景外观的显著影响,使现有卫星影像难以对齐。为此,我们提出MoViX,一种自监督跨视角视频定位框架,可学习视角与季节不变的表示,同时保留对方向的敏感性,以支持精确定位。MoViX采用依赖姿态的正样本采样策略增强方向判别力,并通过时间对齐的困难负样本挖掘,防止模型依赖季节线索产生捷径学习。运动感知的帧采样器选取空间多样化的帧,轻量级时序聚合器强调几何对齐观测,弱化模糊信息。推理阶段,MoViX集成于蒙特卡洛定位框架中,以学习得到的跨视图匹配模块替代手工设计模型。基于熵引导的温度缩放,实现鲁棒的多假设跟踪与高置信度收敛。我们在TartanDrive 2.0数据集上评估,仅用不到30分钟训练数据,测试距离达12.29公里。尽管使用过时卫星影像,MoViX在未见区域中93%的定位误差小于25米,100%小于50米,优于无需环境调优的先进基线。进一步在地理上不同的真实野外数据集(不同机器人平台)上验证了其泛化能力。

原文摘要 · Abstract (English)

Robust cross-view 3-DoF localization in GPS-denied, off-road environments remains challenging due to (1) perceptual ambiguities from repetitive vegetation and unstructured terrain, and (2) seasonal shifts that significantly alter scene appearance, hindering alignment with outdated satellite imagery. To address this, we introduce MoViX, a self-supervised cross-view video localization framework that learns viewpoint- and season-invariant representations while preserving directional awareness essential for accurate localization. MoViX employs a pose-dependent positive sampling strategy to enhance directional discrimination and temporally aligned hard negative mining to discourage shortcut learning from seasonal cues. A motion-informed frame sampler selects spatially diverse frames, and a lightweight temporal aggregator emphasizes geometrically aligned observations while downweighting ambiguous ones. At inference, MoViX runs within a Monte Carlo Localization framework, using a learned cross-view matching module in place of handcrafted models. Entropy-guided temperature scaling enables robust multi-hypothesis tracking and confident convergence under visual ambiguity. We evaluate MoViX on the TartanDrive 2.0 dataset, training on under 30 minutes of data and testing over 12.29 km. Despite outdated satellite imagery, MoViX localizes within 25 meters of ground truth 93% of the time, and within 50 meters 100% of the time in unseen regions, outperforming state-of-the-art baselines without environment-specific tuning. We further demonstrate generalization on a real-world off-road dataset from a geographically distinct site with a different robot platform.

定位自监督野外导航视频匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。