提出无需标定的长距离单目SLAM系统,解决尺度漂移问题。
VGGT-Motion: Motion-Aware Calibration-Free Monocular SLAM for Long-Range Consistency
- 基于光流动态分块,自适应划分子地图并保留转弯结构。
- 采用锚点驱动直接Sim(3)匹配,实现无特征匹配的密集对齐。
- 轻量级子图优化保持全局一致性,支持千米级轨迹运行。
尽管基于3D视觉基础模型的免标定单目SLAM取得进展,长序列下尺度漂移仍严重。传统几何对齐计算成本高,而运动无关分块会破坏上下文连贯性并导致零运动漂移。为此,我们提出VGGT-Motion,一种高效且鲁棒的免标定单目SLAM系统,可实现千米级轨迹的全局一致性。首先,设计运动感知子地图构建机制,利用光流引导自适应分块、剔除静态冗余,并封装转弯以稳定局部几何结构。其次,提出锚点驱动的直接Sim(3)注册策略,通过上下文均衡锚点实现无搜索、像素级稠密对齐及高效回环闭合,无需昂贵特征匹配。最后,采用轻量级子图级位姿图优化,在线性复杂度下维持全局一致性,支持可扩展的长距离运行。实验表明,VGGT-Motion显著提升轨迹精度与效率,在零样本、长距离免标定单目SLAM中达到当前最优性能。
原文摘要 · Abstract (English)
Despite recent progress in calibration-free monocular SLAM via 3D vision foundation models, scale drift remains severe on long sequences. Motion-agnostic partitioning breaks contextual coherence and causes zero-motion drift, while conventional geometric alignment is computationally expensive. To address these issues, we propose VGGT-Motion, a calibration-free SLAM system for efficient and robust global consistency over kilometer-scale trajectories. Specifically, we first propose a motion-aware submap construction mechanism that uses optical flow to guide adaptive partitioning, prune static redundancy, and encapsulate turns for stable local geometry. We then design an anchor-driven direct Sim(3) registration strategy. By exploiting context-balanced anchors, it achieves search-free, pixel-wise dense alignment and efficient loop closure without costly feature matching. Finally, a lightweight submap-level pose graph optimization enforces global consistency with linear complexity, enabling scalable long-range operation. Experiments show that VGGT-Motion markedly improves trajectory accuracy and efficiency, achieving state-of-the-art performance in zero-shot, long-range calibration-free monocular SLAM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。