arXiv:2503.15091cs.ROcs.CV2025-03中稿 · WRC SARA 2024被引 2

用大模型构建分层3D场景图,让机器人更懂室内环境

Intelligent Spatial Perception by Building Hierarchical 3D Scene Graphs for Indoor Scenarios with the Help of LLMs

  • 用大模型自动标注物体与房间等多层节点
  • 融合语义与几何信息,实现精准环境表征
  • 适合需要上下文理解的智能导航任务

本文针对高级智能机器人导航对空间环境全面理解的高需求,提出一种新系统,借助大语言模型(LLMs)构建面向室内场景的分层3D场景图(3DSGs)。该框架包含三层结构:基础层含丰富度量-语义信息,物体层包含精确点云表示和视觉描述符,高层则包括房间、楼层和建筑节点。通过创新应用LLMs,不仅物体节点,连房间等高层节点也能智能准确标注。文中提出基于LLM的房间分类轮询机制,显著提升节点标注的准确性和可靠性。大量数值实验表明,该系统能有效融合语义描述与几何数据,生成准确且全面的环境表征,为上下文感知导航与任务规划提供支持。

原文摘要 · Abstract (English)

This paper addresses the high demand in advanced intelligent robot navigation for a more holistic understanding of spatial environments, by introducing a novel system that harnesses the capabilities of Large Language Models (LLMs) to construct hierarchical 3D Scene Graphs (3DSGs) for indoor scenarios. The proposed framework constructs 3DSGs consisting of a fundamental layer with rich metric-semantic information, an object layer featuring precise point-cloud representation of object nodes as well as visual descriptors, and higher layers of room, floor, and building nodes. Thanks to the innovative application of LLMs, not only object nodes but also nodes of higher layers, e.g., room nodes, are annotated in an intelligent and accurate manner. A polling mechanism for room classification using LLMs is proposed to enhance the accuracy and reliability of the room node annotation. Thorough numerical experiments demonstrate the system's ability to integrate semantic descriptions with geometric data, creating an accurate and comprehensive representation of the environment instrumental for context-aware navigation and task planning.

3D场景图大模型机器人导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。