利用冻结模型隐含线索,实现无IMU全景SLAM高精度定位
Look Up and Look Back: Hidden Attention and Latent Orientation in a Frozen Foundation Model for Panoramic SLAM

- 从冻结模型中间层提取重力方向与跨视图兼容性信号
- 125个真实场景序列全部成功,相对最优基线误差降低30%-88%
- 适合需要轻量化、高鲁棒性的全景导航系统开发者
单目全景SLAM在大角度旋转下受益于丰富的视觉重叠,但仍易受相机倾角、尺度漂移和错误闭环影响。我们发现,一个冻结的全景几何基础模型不仅提供显式几何输出,其内部中间标记还编码了相机坐标系下的重力信息,跨视图注意力则提供了潜在回溯的兼容性线索。基于这些线索,提出HALO-SLAM:重力读出实现无IMU的球面正交化;闭环检测采用三阶段级联策略,结合DBoW2事件级检索、注意力兼容性过滤及对称子图增强的密集几何验证。被接受的回溯可生成局部度量下像素对齐的3D-3D对应关系,从中估计鲁棒的Sim(3)约束,并与序列约束在全局位姿图中联合优化。在五个真实世界全景基准的125个序列上,方法达到100%序列成功率(125/125),且在所有基准上均取得最低绝对轨迹误差(ATE),相较最优的ERP原生基线降低30–88%。
原文摘要 · Abstract (English)
Monocular panoramic SLAM benefits from substantial visual overlap under large camera rotations, yet remains prone to errors caused by camera tilt, scale drift, and false loop closures. We show that a frozen panoramic geometry foundation model provides useful internal cues beyond its explicit geometric outputs: intermediate tokens encode gravity in the camera frame, while cross-view attention provides a compatibility cue for potential revisits. Building on these cues, we present HALO-SLAM. A gravity readout enables IMU-free spherical upright canonicalization. For loop closure, we introduce a cost-aware three-stage cascade combining DBoW2 event-level retrieval, attention-based compatibility filtering, and dense geometric validation through symmetric submap augmentation. Accepted revisits yield pixel-aligned 3D--3D correspondences in both local gauges, from which robust $\mathrm{Sim}(3)$ constraints are estimated and jointly optimized with sequential constraints in a global pose graph. Across 125 sequences from five real-world panoramic benchmarks, our method achieves \textbf{100\%} sequence success (\textbf{125/125}) under the stated criterion and the lowest ATE among the evaluated methods on all five benchmarks, reducing ATE by \textbf{30--88\%} relative to the best ERP-native baseline on each benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。