用大模型生成可扩展的3D语义场景图森林,提升机器人环境理解能力。
From Pixels to Concepts: Growing Rich 3D Semantic Scene Graph Forests utilizing Foundation Models

- 融合视觉语言模型与大语言模型,自动识别并推理抽象概念与关系
- 在uHumans2和ScanNet上验证关系准确性,支持开放词汇任务
- 适用于真实机器人场景,如波士顿动力Spot的物体检索任务
在复杂现实环境中,机器人需在功能语义层面理解周围世界,这要求构建包含几何、语义与关系数据的多层次世界模型。层级化3D场景图通过统一空间框架整合多源信息,解决此问题。然而,现有方法多局限于预定义的关系类别,忽视因果关系或环境上下文等重要语义连接。本文探索利用基础模型构建具有开放语义关系的3D场景图森林,以增强场景理解与机器人任务执行能力。我们提出一种方法:首先由视觉语言模型(VLM)识别实例级概念节点与关系,再由大语言模型(LLM)通过推理推导更广泛、更抽象的概念节点与关系。这些对象节点、概念节点及关系被组装为分层3D场景图森林,并引入概念节点表示抽象概念。在uHumans2和ScanNet室内数据集上评估了生成关系的准确性和相关性。通过使用ScanNet数据和真实室内部署的波士顿动力Spot机器人,验证了场景图森林在下游机器人任务中的适用性。本工作利用基础模型构建更具表现力、语义更深层的3D层次化场景图,展示了其在提升机器人语义与环境理解方面的潜力。
原文摘要 · Abstract (English)
Operating in complex real-world environments requires robots to understand their surroundings on a functional semantic level. This demands a detailed multi-layer world model capturing the complex relations of its surroundings. Hierarchical 3D scene graphs address this challenge by integrating geometric, semantic, and relational data within a unified spatial framework. However, current 3D scene graph approaches often restrict themselves to rigid structures of pre-determined relationship classes, mostly neglecting important semantic connections, like causal connections or environmental contexts. This paper explores the potential of foundation models to build forests of 3D scene graphs with open semantic relationships to improve scene understanding and robotic task execution. We propose a method where instance-specific concept-nodes and relationships are first identified by a VLM and extended upon by a LLM, inferring broader, more abstract concept-nodes and relationships through reasoning. These object-nodes, concept-nodes, and relationships are then assembled into a forest of hierarchical 3D scene graphs, enhanced with concept-nodes to represent abstract concepts. Evaluations were conducted on the uHumans2 and ScanNet indoor dataset, validating the accuracy and relevance of the generated relationships. Downstream suitability of scene-graph forests for robotics applications is demonstrated in an open-vocabulary object-retrieval task utilizing both ScanNet data and a real-world indoor deployment using a Boston Dynamics Spot. This paper leverages foundation models to create more expressive, semantically deep 3D hierarchical scene graphs and demonstrates their potential to advance semantic and environmental understanding in robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。