用单目相机实现多机器人室外高精度协同建图,无需深度传感器。
CoMo3R-SLAM: Collaborative Monocular Dense SLAM with Learned 3D Reconstruction Priors for Outdoor Multi-Agent Systems

- 利用学习的3D重建先验引导前端实时跟踪与局部稠密融合
- 在Tanks and Temples上三项指标最优,Waymo数据集达顶尖水平
- 适合无深度传感器的轻量级多机器人系统部署
多机器人团队在大规模室外环境中实现可扩展、一致的3D感知,依赖于协同稠密SLAM。现有系统通常依赖深度传感器,带来显著的负载、功耗和校准成本。单目RGB相机是轻量化替代方案,但协同单目稠密SLAM仍面临尺度模糊、跨代理数据关联不可靠的问题,尤其在低重叠度、重复结构多的室外场景中,传统特征匹配易失效。为此,我们提出CoMo3R-SLAM,首个基于学习的前馈3D重建先验的室外多智能体协同单目稠密SLAM系统。每个代理运行基于先验引导的前端,实现实时跟踪与局部稠密融合;协调器则执行稠密点云匹配、闭式Sim(3)尺度同步及分段级深度优化的GPU加速全局束调整。系统仅需单目RGB输入,无需深度传感器或参数化内参,即可生成鲁棒的跨代理约束与全局一致的度量地图。在Tanks and Temples和Waymo序列上,CoMo3R-SLAM在四个Tanks and Temples场景中三项指标最优,且在Waymo上表现竞争力,性能达到甚至超越现有RGB-D方法,支持8 FPS在线运行。
原文摘要 · Abstract (English)
Collaborative dense SLAM is essential for multi-robot teams to achieve scalable and consistent 3D perception across large-scale outdoor environments. Existing systems typically depend on depth sensors, incurring significant payload, power, and calibration costs. Monocular RGB cameras are a lightweight alternative, but collaborative monocular dense SLAM remains difficult due to scale ambiguity, unreliable inter-agent data association, especially in outdoor scenes where low overlap and repetitive structures make traditional feature matching unreliable, motivating robust geometric information. We propose CoMo3R-SLAM, the first collaborative monocular dense RGB SLAM system that leverages robust learned feed-forward 3D reconstruction priors for outdoor multi-agent mapping. Each agent runs a prior-guided front-end for real-time tracking and local dense fusion, while a coordinator performs dense pointmap matching for cross-agent verification, closed-form Sim(3) gauge synchronization, and GPU-accelerated global bundle adjustment with segment-level depth optimization. Requiring neither depth sensors nor parametric intrinsics, our system produces robust cross-agent constraints and globally consistent metric maps from monocular RGB alone. On Tanks and Temples and Waymo sequences, CoMo3R-SLAM achieves the best ATE on three of four Tanks and Temples scenes and competitive Waymo accuracy, matching or exceeding state-of-the-art RGB-D methods while running online at 8 FPS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。