arXiv:2604.03377cs.CV2026-04

ViBA通过几何与时间一致性提升视觉匹配鲁棒性,支持持续在线学习。

ViBA: Implicit Bundle Adjustment with Geometric and Temporal Consistency for Robust Visual Matching

  • 隐式可微的捆绑调整框架,联合优化相机位姿与特征点位置。
  • 在EuRoC和UMA数据集上降低12-18%的平移误差,5-10%的旋转误差。
  • 适用于真实场景下需持续学习的视觉导航与定位任务。

现有图像关键点检测与描述方法依赖于带有精确位姿和深度标注的数据集,限制了可扩展性与泛化能力,常导致导航与定位性能下降。我们提出ViBA,一种可持续学习框架,将几何优化与特征学习结合,实现在无约束视频流上的连续在线训练。嵌入标准视觉里程计流程,其包含一个隐式可微的几何残差框架:(i) 初步跟踪网络用于帧间对应关系,(ii) 基于深度的异常值过滤,(iii) 可微全局捆绑调整,通过最小化重投影误差联合优化相机位姿与特征位置。通过结合捆绑调整的几何一致性与跨帧长期时间一致性,ViBA强化了稳定且准确的特征表示。我们在EuRoC和UMA数据集上评估ViBA,相比SuperPoint+SuperGlue、ALIKED、LightGlue等先进方法,在序列上降低12-18%的平均绝对平移误差(ATE)和5-10%的绝对旋转误差(ARE),同时保持实时推理速度(36-91 FPS)。在未见序列上仍保持超过90%的定位精度,证明其出色的泛化能力。结果表明,ViBA支持具有几何与时间一致性的持续在线学习,持续提升真实场景中的导航与定位性能。

原文摘要 · Abstract (English)

Most existing image keypoint detection and description methods rely on datasets with accurate pose and depth annotations, limiting scalability and generalization, and often degrading navigation and localization performance. We propose ViBA, a sustainable learning framework that integrates geometric optimization with feature learning for continuous online training on unconstrained video streams. Embedded in a standard visual odometry pipeline, it consists of an implicitly differentiable geometric residual framework: (i) an initial tracking network for inter-frame correspondences, (ii) depth-based outlier filtering, and (iii) differentiable global bundle adjustment that jointly refines camera poses and feature positions by minimizing reprojection errors. By combining geometric consistency from BA with long-term temporal consistency across frames, ViBA enforces stable and accurate feature representations. We evaluate ViBA on EuRoC and UMA datasets. Compared with state-of-the-art methods such as SuperPoint+SuperGlue, ALIKED, and LightGlue, ViBA reduces mean absolute translation error (ATE) by 12-18% and absolute rotation error (ARE) by 5-10% across sequences, while maintaining real-time inference speeds (FPS 36-91). When evaluated on unseen sequences, it retains over 90% localization accuracy, demonstrating robust generalization. These results show that ViBA supports continuous online learning with geometric and temporal consistency, consistently improving navigation and localization in real-world scenarios.

视觉定位捆绑调整在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。