arXiv:2506.18678cs.CVcs.RO2025-06被引 38

首个支持多智能体协同的神经SLAM框架,解决大场景建图通信难题。

MCN-SLAM: Multi-Agent Collaborative Neural SLAM with Hybrid Implicit Neural Scene Representation

  • 用混合隐式表示与分布式跟踪提升多机协同效率
  • 实现局部与全局一致性,建图精度显著优于现有方法
  • 适合大规模场景实时建图与多智能体系统研究者

神经隐式场景表示在稠密视觉SLAM中展现出良好前景,但现有隐式SLAM算法局限于单智能体,难以处理大场景与长序列。基于NeRF的多智能体框架又受限于通信带宽。为此,我们提出首个分布式多智能体协同神经SLAM框架,结合混合场景表示、分布式相机追踪、跨智能体回环检测及在线知识蒸馏,实现多子地图融合。提出新颖的三平面-网格联合表示方法以提升重建质量;设计跨智能体回环检测机制,保证局部与全局一致性;开发在线蒸馏策略融合子地图信息。此外,目前尚无真实世界中同时提供连续时间轨迹真值与高精度3D网格真值的NeRF/GS基SLAM数据集。因此,我们构建首个涵盖单/多智能体场景、从室内小空间到室外大场景的密集SLAM(DES)数据集,包含高精度3D网格与连续时间相机轨迹真值。实验表明,该方法在建图、跟踪与通信方面均优于现有方法。代码与数据集将开源于https://github.com/dtc111111/mcnslam。

原文摘要 · Abstract (English)

Neural implicit scene representations have recently shown promising results in dense visual SLAM. However, existing implicit SLAM algorithms are constrained to single-agent scenarios, and fall difficulties in large-scale scenes and long sequences. Existing NeRF-based multi-agent SLAM frameworks cannot meet the constraints of communication bandwidth. To this end, we propose the first distributed multi-agent collaborative neural SLAM framework with hybrid scene representation, distributed camera tracking, intra-to-inter loop closure, and online distillation for multiple submap fusion. A novel triplane-grid joint scene representation method is proposed to improve scene reconstruction. A novel intra-to-inter loop closure method is designed to achieve local (single-agent) and global (multi-agent) consistency. We also design a novel online distillation method to fuse the information of different submaps to achieve global consistency. Furthermore, to the best of our knowledge, there is no real-world dataset for NeRF-based/GS-based SLAM that provides both continuous-time trajectories groundtruth and high-accuracy 3D meshes groundtruth. To this end, we propose the first real-world Dense slam (DES) dataset covering both single-agent and multi-agent scenarios, ranging from small rooms to large-scale outdoor scenes, with high-accuracy ground truth for both 3D mesh and continuous-time camera trajectory. This dataset can advance the development of the research in both SLAM, 3D reconstruction, and visual foundation model. Experiments on various datasets demonstrate the superiority of the proposed method in both mapping, tracking, and communication. The dataset and code will open-source on https://github.com/dtc111111/mcnslam.

多智能体神经SLAM隐式表示三维重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。