arXiv:2608.17553cs.RO2026-08

用深度学习提升单目SLAM的尺度一致性,实时精准定位

Scalix: Uncertainty-Aware Scale-Consistent Monocular SLAM

论文配图:Scalix: Uncertainty-Aware Scale-Consistent Monocular SLAM
图 1 · 摘自论文原文
  • 将像素级深度不确定性和帧级尺度不确定性融入概率因子图优化
  • 在室内外大场景中实现优于现有方法的度量尺度精度和实时性
  • 适合对定位精度要求高、依赖单目相机的机器人应用

摄像头因体积小、视觉信息丰富,广泛应用于机器人。单目SLAM虽能以最少配置理解环境,但存在固有的尺度模糊问题。常见解决方案是融合多模态传感器(如视觉-惯性系统),但在匀速运动时仍无法解算尺度,而深度学习生成的深度图常存在噪声和帧间尺度不一致。本文提出Scalix,一种实时单目SLAM框架,通过将学习到的深度线索与概率因子图结合,实现度量尺度状态估计。其创新在于为现有单目深度模型引入像素级深度不确定性和帧级尺度不确定性,将尺度预测视为独立观测量参与优化,利用多视角数据关联提升尺度一致性。在大型室内外环境中实验表明,Scalix在度量尺度与非度量尺度基准上均达到领先性能,同时保持实时运行与良好泛化能力。

原文摘要 · Abstract (English)

Cameras are ubiquitous sensors in robotics due to their compact form factor and the perceptual richness captured through visual information. Monocular SLAM enables robots to understand the environment with a minimum setup, however, it inherently suffers from scale ambiguity. A common solution is to provide multi-modal sensor configurations, such as visual-inertial systems, where scale is observable unless the robot navigates under a constant-velocity motion, a common scenario in mobile robotics. With the advent of deep-learning, geometric foundation models have been used to address this problem, but the depths maps are often noisy and scale-inconsistent across frames. In this paper, we propose Scalix, a real-time monocular SLAM framework that achieves metric-scale state estimation by integrating learned depth cues into a probabilistic factor-graph formulation. By augmenting existing monocular depth models with both per-pixel depth uncertainty and per-frame scale uncertainty, Scalix treats scale predictions as independent measurements within its optimization, leading to improved scale consistency through multi-view data associations. Experiments in large-scale outdoor and indoor environments demonstrate state-of-the-art performance on both metric and up-to-scale benchmarks while maintaining real-time operation and generalization.

单目SLAM深度估计概率建模机器人定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。