arXiv:2601.19887cs.CVcs.RO2026-01被引 25

VGGT-SLAM 2.0实现高精度实时三维重建,提升定位准确率并支持复杂场景下闭环检测。

VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction

  • 设计新因子图消除15自由度漂移和平面退化问题。
  • 利用VGGT注意力层免费实现图像匹配验证,闭环成功率提升。
  • 在真实机器人上实现实时运行,跨场景测试误差比原系统低23%。

我们提出VGGT-SLAM 2.0,一种实时彩色图像前馈式SLAM系统,显著改进了基于VGGT构建子地图的增量对齐效果。首先,通过全新因子图设计,消除了原系统中高维15自由度漂移和平面退化问题,同时仍能处理相机内参未知下的重建模糊性。其次,通过对VGGT注意力层的研究发现,其中一层可无需额外训练即可用于图像检索验证,有效拒绝误匹配并完成更多闭环。最后,实验表明该系统可轻松适配开放集目标检测,并在搭载Jetson Thor的地面机器人上实现在线实时运行。测试环境涵盖杂乱室内公寓、办公室及4,200平方英尺谷仓;在TUM数据集上,其姿态误差比VGGT-SLAM降低约23%,达到最高精度。代码将在发表后公开。

原文摘要 · Abstract (English)

We present VGGT-SLAM 2.0, a real-time RGB feed-forward SLAM system which substantially improves upon VGGT-SLAM for incrementally aligning submaps created from VGGT. Firstly, we remove high-dimensional 15-degree-of-freedom drift and planar degeneracy from VGGT-SLAM by creating a new factor graph design while still addressing the reconstruction ambiguity of VGGT given unknown camera intrinsics. Secondly, by studying the attention layers of VGGT, we show that one of the layers is well suited to assist in image retrieval verification for free without additional training, which enables both rejecting false positive matches and allows for completing more loop closures. Finally, we conduct a suite of experiments which includes showing VGGT-SLAM 2.0 can easily be adapted for open-set object detection and demonstrating real-time performance while running online onboard a ground robot using a Jetson Thor. We test in environments ranging from cluttered indoor apartments and office scenes to a 4,200 square foot barn, and we also demonstrate VGGT-SLAM 2.0 achieves the highest accuracy on the TUM dataset with about 23 percent less pose error than VGGT-SLAM. Code will be released upon publication.

SLAM三维重建实时系统视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。