arXiv:2607.05348cs.CVcs.RO2026-07

用3D场景图提升开放词汇3D理解,捕捉物体间关系增强语义推理。

Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis

论文配图:Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis
图 1 · 摘自论文原文
  • 构建多视角3D场景图,用视觉语言模型推断物体关系并剔除不合理连接。
  • 在ScanNetV2等数据集上达到新高,实例级一致性与类别判别力显著提升。
  • 适合关注3D场景理解、关系建模及跨模态推理的研究者。

开放词汇3D场景理解旨在通过迁移视觉-语言模型的语义知识,对3D场景进行超出预定义类别的分割。现有方法虽将语言对齐的2D特征映射至3D,但多依赖无上下文的语义表示,忽视了物体关系在上下文优化中的作用。本文提出RelGraphOV,一种基于关系感知的框架,利用3D场景图增强开放词汇3D理解。该方法通过多视角观测构建关系场景图,借助视觉-语言推理推断物体间关系,并剔除几何上不合理的连接,无需人工标注关系。为聚合关系上下文同时避免特征干扰,引入自适应门控双流上下文GAT,分离密集几何特征与语义CLIP嵌入,进行边引导的消息传递,并自适应融合互补语义。层次化对比目标进一步促进实例级一致性与类别级判别力。在ScanNetV2、ScanNet200、ScanNet++和Replica上的实验表明,该方法具有优异性能与泛化能力。

原文摘要 · Abstract (English)

Open-vocabulary 3D scene understanding aims to segment 3D scenes beyond predefined categories by transferring semantic knowledge from vision-language models. Existing methods have advanced this task by lifting language-aligned 2D features into 3D, yet they often rely on context-independent semantic representations, leaving object relationships underexplored for contextual refinement. We propose RelGraphOV, a relationship-aware framework that uses 3D scene graphs to enhance open-vocabulary 3D understanding. Our method constructs relational scene graphs from multi-view observations by leveraging vision-language reasoning to infer object relationships and prune geometrically implausible connections, without manual relationship annotations. To aggregate relational context while avoiding feature interference, we introduce an Adaptive Gated Dual-Stream Contextual GAT that separates dense geometric features and semantic CLIP embeddings, performs edge-guided message passing, and adaptively fuses complementary semantics. A hierarchical contrastive objective further promotes instance-level consistency and category-level discrimination. Experiments on ScanNetV2, ScanNet200, ScanNet$++$, and Replica demonstrate strong performance and generalization ability. Project Page: https://cxavireh.github.io/relgraphov-projectpage

3D场景理解关系建模开放词汇场景图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。