arXiv:2512.25008cs.CV2025-12AAAI

用深度基础模型提升单目稠密SLAM的精度与鲁棒性

FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM

  • 通过融合深度模型引导光流,实现几何一致的匹配与定位
  • 在多个数据集上达到更优轨迹精度和稠密重建质量,实时运行18帧/秒
  • 自适应调整不确定区域的更新策略,适合复杂场景下的实际应用

我们提出FoundationSLAM,一种基于学习的单目稠密视觉SLAM系统,解决以往基于光流方法中缺乏几何一致性的问题。核心思想是利用深度基础模型的引导,将光流估计与几何推理相结合。为此,我们设计了混合光流网络,生成具有几何感知的对应关系,实现跨关键帧的一致深度与位姿推断。为保证全局一致性,提出双向一致性束调整层,在多视角约束下联合优化关键帧位姿与深度。此外,引入可靠性感知精化机制,通过区分可靠与不确定区域,动态调节光流更新过程,形成匹配与优化间的闭环反馈。大量实验表明,FoundationSLAM在多个挑战性数据集上实现更优的轨迹精度与稠密重建效果,同时以18 FPS实现实时运行,展现出强泛化能力与实际应用价值。

原文摘要 · Abstract (English)

We present FoundationSLAM, a learning-based monocular dense SLAM system that addresses the absence of geometric consistency in previous flow-based approaches for accurate and robust tracking and mapping. Our core idea is to bridge flow estimation with geometric reasoning by leveraging the guidance from foundation depth models. To this end, we first develop a Hybrid Flow Network that produces geometry-aware correspondences, enabling consistent depth and pose inference across diverse keyframes. To enforce global consistency, we propose a Bi-Consistent Bundle Adjustment Layer that jointly optimizes keyframe pose and depth under multi-view constraints. Furthermore, we introduce a Reliability-Aware Refinement mechanism that dynamically adapts the flow update process by distinguishing between reliable and uncertain regions, forming a closed feedback loop between matching and optimization. Extensive experiments demonstrate that FoundationSLAM achieves superior trajectory accuracy and dense reconstruction quality across multiple challenging datasets, while running in real-time at 18 FPS, demonstrating strong generalization to various scenarios and practical applicability of our method.

SLAM深度学习稠密重建实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。