arXiv:2409.12518cs.ROcs.AI2024-09ICRA被引 27

Hier-SLAM用分层语义表示,实现高精度3D语义建图与超大规模场景扩展。

Hier-SLAM: Scaling-up Semantics in SLAM with a Hierarchically Categorical Gaussian Splatting

  • 采用分层语义编码,压缩参数量并提升环境理解能力
  • 支持500+语义类别,渲染帧率达2000 FPS,训练时间显著减少
  • 适合需要高效、大规模语义建图的自动驾驶与机器人应用

我们提出 Hier-SLAM,一种基于分层分类高斯点阵的语义3D高斯溅射SLAM方法,具备新颖的层次化语义表示,可实现高精度全局3D语义映射、可扩展性及显式3D世界语义标签预测。传统语义SLAM系统在复杂环境下的参数量急剧增长,导致场景理解成本高昂。为此,我们引入一种新型层次化表示,将语义信息以紧凑形式嵌入3D高斯溅射,并利用大语言模型(LLMs)增强表达能力。同时设计了一种新型语义损失,通过层级内与跨层级联合优化来提升层次化语义信息。进一步优化整个SLAM系统,显著提升跟踪与建图性能。实验表明,该方法在建图与跟踪精度上优于现有密集型SLAM方法,且运算速度提升2倍。在语义渲染效果与现有方法相当的前提下,大幅降低存储与训练时间开销。带语义信息的渲染帧率达2000 FPS,无语义时达3000 FPS。最显著的是,能处理包含超过500个语义类别的复杂真实场景,凸显其卓越的可扩展能力。开源代码已发布于 https://github.com/LeeBY68/Hier-SLAM。

原文摘要 · Abstract (English)

We propose Hier-SLAM, a semantic 3D Gaussian Splatting SLAM method featuring a novel hierarchical categorical representation, which enables accurate global 3D semantic mapping, scaling-up capability, and explicit semantic label prediction in the 3D world. The parameter usage in semantic SLAM systems increases significantly with the growing complexity of the environment, making it particularly challenging and costly for scene understanding. To address this problem, we introduce a novel hierarchical representation that encodes semantic information in a compact form into 3D Gaussian Splatting, leveraging the capabilities of large language models (LLMs). We further introduce a novel semantic loss designed to optimize hierarchical semantic information through both inter-level and cross-level optimization. Furthermore, we enhance the whole SLAM system, resulting in improved tracking and mapping performance. Our \MethodName{} outperforms existing dense SLAM methods in both mapping and tracking accuracy, while achieving a 2x operation speed-up. Additionally, it achieves on-par semantic rendering performance compared to existing methods while significantly reducing storage and training time requirements. Rendering FPS impressively reaches 2,000 with semantic information and 3,000 without it. Most notably, it showcases the capability of handling the complex real-world scene with more than 500 semantic classes, highlighting its valuable scaling-up capability. The open-source code is available at https://github.com/LeeBY68/Hier-SLAM

3D建图语义分割高斯溅射可扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。