提出可识别部件的3D场景图生成框架,提升真实环境理解能力
OP3DSG: Open-Vocabulary Part-Aware 3D Scene Graph Generation for Real-World Environments

- 基于部件感知检测与融合,保留小而重要的交互组件
- 通过几何初始化图结构+大模型优化,减少错误关系预测
- 适用于机器人任务的细粒度环境理解,支持多层级推理
3D场景图(3DSG)为3D环境提供了紧凑且结构化的抽象。尽管基础模型的发展推动了开放词汇3DSG生成,但现有方法仍以物体为中心,关系信息编码有限,难以满足真实场景中对细粒度理解的需求。本文提出OP3DSG,一种开放词汇、部件感知的3DSG生成框架,能统一建模物体、交互部件、空间关系、功能关系与可操作性。该框架结合物体-部件知识引导的检测与部件感知的3D融合,保留微小且交互相关的组件;采用几何初始化先验图与LLM驱动的优化,减少虚假关系预测,同时实现高效图构建。为系统评估统一3DSG构建,我们引入UniGraph3D基准,专用于部件感知与多层级关系推理。实验表明,OP3DSG达到当前最优性能,并在多样化真实世界机器人任务中展现出作为感知骨干的有效性。
原文摘要 · Abstract (English)
3D scene graphs (3DSGs) provide a compact and structured abstraction of 3D environments. Although advances in foundation models have enabled open-vocabulary 3DSG generation, existing approaches remain object-centric and encode limited relational information -- restricting their applicability in real-world scenarios that require fine-grained understanding. We propose OP3DSG, an open-vocabulary part-aware 3DSG generation framework that constructs unified graphs that jointly model objects, interactive parts, spatial relations, functional relations, and affordances. OP3DSG integrates object-part knowledge-guided detection with part-aware 3D fusion to preserve small and interaction-relevant components, and employs a geometry-initialized prior graph with LLM-based refinement to reduce spurious relational predictions while enabling efficient graph construction. To systematically evaluate unified 3D scene graph construction, we introduce UniGraph3D, a benchmark designed for part-aware perception and multi-level relational reasoning. Experimental results show that OP3DSG achieves state-of-the-art performance and demonstrates its effectiveness as a perception backbone in diverse real-world robotics tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。