通过学习短时一致性先验,提升复杂环境下RGB-D里程计的鲁棒性。
Con-DSO: Learning Short-Horizon Consistency Priors for RGB-D Direct Sparse Odometry

- 用相邻帧预测像素级光度与深度几何不确定性
- 在ICL-NUIM上轨迹误差降低超20%,其他数据集降50%-80%
- 适合动态场景、光照变化等复杂环境下的机器人定位
视觉里程计是机器人和增强现实的基础。基于RGB-D的直接稀疏里程计虽利用了度量深度,但在动态物体、遮挡、光照变化和不可靠深度等挑战下性能下降,因直接对齐依赖短时光度和深度几何一致性假设。现有方法通过语义过滤、显式遮挡推理、光照适应或手工几何准则缓解问题,但常需外部模块或固定假设,难以统一应对多样挑战。本文提出Con-DSO,一种一致性感知的RGB-D直接稀疏里程计框架,通过时间相邻帧对预测密集光度与深度几何一致性不确定性。一致性网络基于光流引导的光度误差和投影深度一致性误差训练,使一致性违反以像素级不确定性表征。这些成对不确定性被转化为关键帧跟踪的质量先验,并通过质量感知支持像素选择和解耦光度-几何加权实现姿态估计,实现对不可靠观测的连续衰减而非硬拒绝或阈值门控。五个公开RGB-D基准测试显示,相比直接基线有显著提升,在ICL-NUIM上轨迹误差绝对降低超过20%,在RGB-D Scenes V2、TUM/Bonn Dynamic和OpenLORIS序列上降低50%–80%。
原文摘要 · Abstract (English)
Visual odometry (VO) is a fundamental component in robotics and augmented reality. RGB-D direct VO benefits from metric depth measurements, but it can degrade in challenging environments, where dynamic objects, occlusions, illumination changes, and unreliable depth violate the short-horizon photometric and depth-geometric consistency assumptions used by direct alignment. Existing approaches mitigate these issues through semantic filtering, explicit occlusion reasoning, illumination adaptation, or hand-crafted geometric criteria, but often rely on external modules or fixed assumptions tailored to individual failure modes, limiting their flexibility and ability to handle diverse challenges in a unified manner. In this work, we propose Con-DSO, a consistency-aware RGB-D direct sparse odometry framework that predicts dense photometric and depth-geometric consistency uncertainty from temporally adjacent RGB-D frame pairs. The consistency network is trained using flow-guided photometric errors and projective depth-consistency errors, allowing consistency violations to be represented as pixel-level uncertainty. These pairwise uncertainty predictions are converted into a host-side quality prior for keyframe-based tracking. The prior is then applied to VO through quality-aware support-pixel selection and decoupled photometric-geometric weighting during pose estimation, enabling continuous attenuation of unreliable observations rather than hard rejection or threshold-based gating. Experiments on five public RGB-D benchmarks show substantial gains over direct RGB-D VO baselines, with over 20\% absolute trajectory error reduction on ICL-NUIM and 50\%--80\% reductions on RGB-D Scenes V2, TUM/Bonn Dynamic, and OpenLORIS sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。