arXiv:2509.21602cs.RO2025-09被引 2

用大模型提供物体尺寸方向先验,提升室内物体级SLAM精度

Real-Time Indoor Object SLAM with LLM-Enhanced Priors

  • 用大模型生成物体尺寸与朝向的先验知识,融入图优化SLAM框架
  • 在TUM RGB-D和3RScan上比最新基线地图精度提升36.8%
  • 适合需要高精度语义建图的机器人导航与增强现实场景

物体级同步定位与地图构建(SLAM)通过引入语义信息实现高层场景理解,但稀疏观测导致优化过程欠约束。以往工作使用常识知识增加约束,但获取这些先验需大量人工且泛化性差。本文利用大语言模型(LLM)生成物体几何属性(如尺寸、朝向)的常识知识,作为图优化框架中的先验因子。该方法在物体观测稀疏的初始阶段尤为有效。我们构建了完整系统,实现了对稀疏物体特征的鲁棒数据关联,支持实时物体级SLAM。在TUM RGB-D和3RScan数据集上的实验表明,相比最新基线,地图精度提升36.8%。补充视频展示了真实世界中的实时性能表现。

原文摘要 · Abstract (English)

Object-level Simultaneous Localization and Mapping (SLAM), which incorporates semantic information for high-level scene understanding, faces challenges of under-constrained optimization due to sparse observations. Prior work has introduced additional constraints using commonsense knowledge, but obtaining such priors has traditionally been labor-intensive and lacks generalizability across diverse object categories. We address this limitation by leveraging large language models (LLMs) to provide commonsense knowledge of object geometric attributes, specifically size and orientation, as prior factors in a graph-based SLAM framework. These priors are particularly beneficial during the initial phase when object observations are limited. We implement a complete pipeline integrating these priors, achieving robust data association on sparse object-level features and enabling real-time object SLAM. Our system, evaluated on the TUM RGB-D and 3RScan datasets, improves mapping accuracy by 36.8\% over the latest baseline. Additionally, we present real-world experiments in the supplementary video, demonstrating its real-time performance.

SLAM大模型语义建图实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。