arXiv:2601.16020cs.CVcs.RO2026-01被引 3

用强化学习自动选关键帧,让视觉里程计更准更快

Keyframe-Based Feed-Forward Visual Odometry

  • 用强化学习自动决定何时提取关键帧,不依赖人工规则
  • 在多个真实数据集上,精度和速度均优于现有前沿方法
  • 适合追求高效高精度的自动驾驶与机器人定位场景

视觉基础模型的兴起彻底改变了视觉里程计(VO)与SLAM,使单次前馈网络即可完成位姿估计与稠密重建。然而,当前基于基础模型的方法(如VGGT-Long)通常无差别处理原始图像序列,导致计算冗余和性能下降,尤其在帧间视差小、立体信息不足时表现更差。将传统几何启发式方法融入其中极具挑战,因这些方法依赖高维潜在表示而非显式几何度量。为此,我们提出一种新型关键帧驱动的前馈式视觉里程计。不同于手工设计规则,我们的方法通过强化学习以数据驱动方式生成自适应关键帧策略,使其与底层基础模型的内在特性相匹配。我们在TartanAir数据集上训练智能体,并在多个真实世界数据集上进行广泛评估。实验结果表明,该方法在各项指标上均显著优于现有最先进前馈式VO方法。

原文摘要 · Abstract (English)

The emergence of visual foundation models has revolutionized visual odometry~(VO) and SLAM, enabling pose estimation and dense reconstruction within a single feed-forward network. However, unlike traditional pipelines that leverage keyframe methods to enhance efficiency and accuracy, current foundation model based methods, such as VGGT-Long, typically process raw image sequences indiscriminately. This leads to computational redundancy and degraded performance caused by low inter-frame parallax, which provides limited contextual stereo information. Integrating traditional geometric heuristics into these methods is non-trivial, as their performance depends on high-dimensional latent representations rather than explicit geometric metrics. To bridge this gap, we propose a novel keyframe-based feed-forward VO. Instead of relying on hand-crafted rules, our approach employs reinforcement learning to derive an adaptive keyframe policy in a data-driven manner, aligning selection with the intrinsic characteristics of the underlying foundation model. We train our agent on TartanAir dataset and conduct extensive evaluations across several real-world datasets. Experimental results demonstrate that the proposed method achieves consistent and substantial improvements over state-of-the-art feed-forward VO methods.

视觉里程计强化学习关键帧前馈网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。