用场景图做地图底层架构,让机器人理解环境更像人。
A Scene Graph Backed Approach to Open Set Semantic Mapping
- 以3D场景图为核心构建地图,实时更新结构
- 支持大场景长期运行,保持拓扑一致性和计算效率
- 适合需要可解释、可信推理的智能体系统
尽管开放集语义映射和三维语义场景图(3DSSG)是机器人感知中的成熟范式,但在大规模真实环境中有效部署以支持高层推理仍面临重大挑战。现有方法通常将感知与表示分离,将场景图作为事后生成的衍生层,限制了一致性与可扩展性。本文提出一种以3DSSG为底层基础的地图架构,将其作为整个建图过程的核心知识表示。通过借鉴增量场景图预测的前期工作,我们在环境探索过程中实时推断并更新图结构,确保在长时间大尺度环境下地图仍保持拓扑一致性与计算高效性。通过维护一个显式的、空间定位的表示,支持扁平与分层拓扑结构,我们弥合了子符号传感器数据与高层符号推理之间的鸿沟。由此形成的稳定、可验证结构,使知识驱动框架(如知识图谱、本体、大语言模型)可直接利用,提升智能体的可解释性、可信度及与人类概念的一致性。
原文摘要 · Abstract (English)
While Open Set Semantic Mapping and 3D Semantic Scene Graphs (3DSSGs) are established paradigms in robotic perception, deploying them effectively to support high-level reasoning in large-scale, real-world environments remains a significant challenge. Most existing approaches decouple perception from representation, treating the scene graph as a derivative layer generated post hoc. This limits both consistency and scalability. In contrast, we propose a mapping architecture where the 3DSSG serves as the foundational backend, acting as the primary knowledge representation for the entire mapping process. Our approach leverages prior work on incremental scene graph prediction to infer and update the graph structure in real-time as the environment is explored. This ensures that the map remains topologically consistent and computationally efficient, even during extended operations in large-scale settings. By maintaining an explicit, spatially grounded representation that supports both flat and hierarchical topologies, we bridge the gap between sub-symbolic raw sensor data and high-level symbolic reasoning. Consequently, this provides a stable, verifiable structure that knowledge-driven frameworks, ranging from knowledge graphs and ontologies to Large Language Models (LLMs), can directly exploit, enabling agents to operate with enhanced interpretability, trustworthiness, and alignment to human concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。