arXiv:2505.12549cs.CV2025-05NeurIPS被引 155

用SL(4)流形优化,解决单目相机无标定下的稠密重建难题。

VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold

论文配图:VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
图 1 · 摘自论文原文
  • 在SL(4)流形上优化子图间15自由度的投影变换
  • 实测在长视频序列下地图质量显著优于原方法
  • 适合高精度稠密重建与无标定单目系统应用

我们提出VGGT-SLAM,一种基于前馈场景重建方法VGGT的增量式、全局对齐的稠密单目RGB SLAM系统。现有方法使用相似变换(平移、旋转、缩放)对齐子图,但在无标定相机下效果不佳。我们重新审视重建模糊性:在无相机运动或场景结构假设的情况下,场景仅可恢复至15自由度的射影变换。由此启发,我们在SL(4)流形上优化,估计相邻子图间的15自由度同构变换,并纳入潜在回环约束。大量实验验证,相比原方法因高GPU需求无法处理的长视频序列,VGGT-SLAM实现了更优的地图质量。

原文摘要 · Abstract (English)

We present VGGT-SLAM, a dense RGB SLAM system constructed by incrementally and globally aligning submaps created from the feed-forward scene reconstruction approach VGGT using only uncalibrated monocular cameras. While related works align submaps using similarity transforms (i.e., translation, rotation, and scale), we show that such approaches are inadequate in the case of uncalibrated cameras. In particular, we revisit the idea of reconstruction ambiguity, where given a set of uncalibrated cameras with no assumption on the camera motion or scene structure, the scene can only be reconstructed up to a 15-degrees-of-freedom projective transformation of the true geometry. This inspires us to recover a consistent scene reconstruction across submaps by optimizing over the SL(4) manifold, thus estimating 15-degrees-of-freedom homography transforms between sequential submaps while accounting for potential loop closure constraints. As verified by extensive experiments, we demonstrate that VGGT-SLAM achieves improved map quality using long video sequences that are infeasible for VGGT due to its high GPU requirements.

SLAM稠密重建单目流形优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。