用几何验证提升视觉测量,让视觉惯性里程计更抗分布偏移。
Does Robust VIO Need More Learning? Geometry-Verified Visual Measurements under Distribution Shift

- 只在视觉测量环节用学习,其余步骤保持显式几何建模。
- 在欧罗巴、户外等复杂场景下定位误差低于0.1%且稳定。
- 适合需要高鲁棒性的自动驾驶与机器人导航场景。
视觉惯性里程计(VIO)中引入学习已成趋势,但训练分布外的部署条件常导致性能下降。本文提出一种极简学习的立体VIO框架:仅用SEA-RAFT生成稠密立体匹配和不确定性预测,其余如时序跟踪、几何验证和状态估计均保持显式。稠密光流在稀疏特征点采样,经不确定性与立体对极一致性过滤后,以加权重投影因子融入滑动窗口立体-惯性估计算法。该不确定性进一步传播至3D高斯映射中的各向异性建模。在EuRoC、VIODE和4Seasons数据集上,该方法在运动模糊、动态场景、光照变化及室内外大分布偏移下均实现稳定且精确的估计。消融实验证明,仅依赖学习匹配无法提升鲁棒性,关键在于学习与几何验证、不确定性加权的协同。结果表明,在分布外条件下,精巧整合的学习视觉测量比深度学习整条流水线更有效。代码与配置将在录用后开源。
原文摘要 · Abstract (English)
Learning is increasingly introduced into visual-inertial odometry (VIO), ranging from learned feature front-ends to learning-dominant motion and geometry estimation. However, learning more of the pipeline does not necessarily improve robustness when deployment conditions differ from the training distribution. This work asks whether robust VIO under distribution shift truly requires deeper learned estimation, or whether learning can be confined to visual measurement generation. We propose a minimal-learning stereo VIO framework in which SEA-RAFT is used only to propose dense stereo correspondences and predict their uncertainty, while temporal tracking, geometric verification, and state estimation remain explicit. Dense flow is sampled at sparse feature locations, filtered using predicted uncertainty and stereo epipolar consistency, and incorporated into a sliding-window stereo-inertial estimator through uncertainty-weighted reprojection factors. The same uncertainty is further propagated through stereo triangulation for downstream anisotropic 3D Gaussian mapping. Experiments on EuRoC, VIODE, and 4Seasons demonstrate accurate and stable estimation under motion blur, dynamic scenes, illumination changes, and large indoor-to-outdoor distribution shifts. Ablations show that learned flow alone is insufficient: the gains arise from combining learned correspondence proposals with geometric verification and uncertainty-aware weighting. These results suggest that, for OOD-robust VIO, carefully integrated learned visual measurements can be more effective than learning a larger fraction of the estimation pipeline. Code and configs for the benchmark will be open-source upon acceptance. A supplementary video is available at https://drive.google.com/file/d/1EVRhOkhanmNXHbQS1Vr80FoEIAYOYOV2/view
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。