让3D场景理解突破固定标签限制,支持语言查询与动态推理。
Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
- 用视觉语言模型+检索增强,实现无固定标签的动态场景图生成。
- 在3DSSG和Replica上四项任务表现优于传统方法,跨环境泛化强。
- 适合做智能机器人、AR/VR交互的开发者和研究者参考。
开放世界中的3D场景理解对视觉与机器人技术构成根本挑战,主要源于封闭词汇监督和静态标注的局限性。为此,我们提出一种融合检索增强推理的开放世界3D场景图生成统一框架,实现可泛化且交互式的3D场景理解。该方法将视觉语言模型(VLMs)与基于检索的推理相结合,支持多模态探索与语言引导交互。框架包含两个核心组件:(1) 动态场景图生成模块,无需固定标签集即可检测物体并推断语义关系;(2) 检索增强推理管道,将场景图编码至向量数据库,支持文本/图像条件查询。我们在3DSSG与Replica基准上针对四项任务——场景问答、视觉定位、实例检索与任务规划——进行评估,展现出卓越的泛化能力与性能。结果表明,将开放词汇感知与检索增强推理结合,是实现可扩展3D场景理解的有效路径。
原文摘要 · Abstract (English)
Understanding 3D scenes in open-world settings poses fundamental challenges for vision and robotics, particularly due to the limitations of closed-vocabulary supervision and static annotations. To address this, we propose a unified framework for Open-World 3D Scene Graph Generation with Retrieval-Augmented Reasoning, which enables generalizable and interactive 3D scene understanding. Our method integrates Vision-Language Models (VLMs) with retrieval-based reasoning to support multimodal exploration and language-guided interaction. The framework comprises two key components: (1) a dynamic scene graph generation module that detects objects and infers semantic relationships without fixed label sets, and (2) a retrieval-augmented reasoning pipeline that encodes scene graphs into a vector database to support text/image-conditioned queries. We evaluate our method on 3DSSG and Replica benchmarks across four tasks-scene question answering, visual grounding, instance retrieval, and task planning-demonstrating robust generalization and superior performance in diverse environments. Our results highlight the effectiveness of combining open-vocabulary perception with retrieval-based reasoning for scalable 3D scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。