arXiv:2507.00243cs.CV2025-07中稿 · WACV 2026被引 1

用对比学习重构视觉里程计,让模型更可解释、更通用。

VOCAL: Visual Odometry via ContrAstive Learning

  • 将视觉里程计转为特征排序问题,通过对比学习对齐相机状态
  • 在KITTI数据集上实现更高可解释性,兼容多模态数据
  • 适合关注模型透明度与通用性的自动驾驶研究者

视觉里程计(VO)的突破性进展彻底重塑了机器人领域,实现了高精度的相机状态估计,这对现代自主系统至关重要。尽管如此,许多基于学习的VO方法依赖于严格的几何假设,在可解释性方面表现不足,且缺乏完全数据驱动框架下的理论基础。为此,我们提出VOCAL(基于对比学习的视觉里程计),将VO重新定义为一个标签排序任务。通过融合贝叶斯推理与表示学习框架,VOCAL使视觉特征能够反映相机状态。其排序机制促使相似相机状态在隐空间中收敛为一致且空间连贯的表示。这种策略不仅提升了学习特征的可解释性,还确保了与多模态数据源的兼容性。在KITTI数据集上的大量评估表明,VOCAL显著增强了可解释性与灵活性,推动视觉里程计向更通用、可解释的空间智能迈进。

原文摘要 · Abstract (English)

Breakthroughs in visual odometry (VO) have fundamentally reshaped the landscape of robotics, enabling ultra-precise camera state estimation that is crucial for modern autonomous systems. Despite these advances, many learning-based VO techniques rely on rigid geometric assumptions, which often fall short in interpretability and lack a solid theoretical basis within fully data-driven frameworks. To overcome these limitations, we introduce VOCAL (Visual Odometry via ContrAstive Learning), a novel framework that reimagines VO as a label ranking challenge. By integrating Bayesian inference with a representation learning framework, VOCAL organizes visual features to mirror camera states. The ranking mechanism compels similar camera states to converge into consistent and spatially coherent representations within the latent space. This strategic alignment not only bolsters the interpretability of the learned features but also ensures compatibility with multimodal data sources. Extensive evaluations on the KITTI dataset highlight VOCAL's enhanced interpretability and flexibility, pushing VO toward more general and explainable spatial intelligence.

视觉里程计对比学习可解释性机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。