arXiv:2603.05097cs.RO2026-03被引 2

用自适应多视角关键帧选择提升单目稠密建图精度

AIM-SLAM: Dense Monocular SLAM via Adaptive and Informative Multi-View Keyframe Prioritization with Foundation Model

  • 基于视觉几何变换器预测点云,动态筛选最优多视角关键帧
  • 在真实数据集上实现领先位姿估计与高精度稠密重建
  • 适合需要高效稠密建图的机器人与AR应用

近期几何基础模型为解决单目视觉同时定位与地图构建(SLAM)中的稠密重建挑战提供了新思路。尽管几何基础模型支持可变输入视角,但现有方法仍局限于双视图对或固定长度输入,缺乏对几何上下文的充分考量。为此,我们提出AIM-SLAM,一种利用视觉几何接地变压器(VGGT)生成稠密点云并结合自适应、信息感知的多视角关键帧优先级机制的稠密单目SLAM框架。具体地,我们设计了选择性信息与几何感知多视角适配(SIGMA)模块,通过体素重叠和信息增益筛选候选关键帧,并自适应确定其数量。此外,我们提出联合多视角Sim(3)优化,强制所选视角间保持一致对齐,显著提升位姿估计精度。在真实世界数据集上的实验表明,该系统在位姿估计性能和稠密重建准确性方面均达到当前最优水平。系统支持ROS集成,代码已开源于https://aimslam.github.io/。

原文摘要 · Abstract (English)

Recent advances in geometric foundation models have emerged as a promising alternative for addressing the challenge of dense reconstruction in monocular visual simultaneous localization and mapping (SLAM). Although geometric foundation models enable SLAM to leverage variable input views, the previous methods remain confined to two-view pairs or fixed-length inputs without sufficient deliberation of geometric context for view selection. To tackle this problem, we propose AIM-SLAM, a dense monocular SLAM framework that exploits an adaptive and informative multi-view keyframe prioritization with dense pointmap predictions from visual geometry grounded transformer (VGGT). Specifically, we introduce the selective information- and geometric-aware multi-view adaptation (SIGMA) module, which employs voxel overlap and information gain to retrieve a candidate set of keyframes and adaptively determine its size. Furthermore, we formulate a joint multi-view Sim(3) optimization that enforces consistent alignment across selected views, substantially improving pose estimation accuracy. The effectiveness of AIM-SLAM is demonstrated on real-world datasets, where it achieves state-of-the-art pose estimation performance and accurate dense reconstruction results. Our system supports ROS integration, with code is available at https://aimslam.github.io/.

单目SLAM稠密重建基础模型关键帧选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。