arXiv:2605.15753cs.ROcs.CV2026-05

构建可理解家具功能关系的3D场景图,支持复杂室内环境推理。

Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces

论文配图:Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
图 1 · 摘自论文原文
  • 基于2D视觉定位与3D图优化,实现细粒度功能边推断
  • 在动态视角下仍能准确关联跨帧物体,错误率降低37%
  • 支持多层次功能结构重建,适用于机器人交互场景

功能性3D场景图提供了灵活的3D场景理解与机器人操作表示,由物体节点、交互元素和功能关系边构成。然而,现有基准覆盖有限且流程设计过于简单,主要关注大规模家具而缺乏层次结构,限制了其潜力。为此,本文引入密集桌面物体和显式多层级功能关系,扩展基准覆盖范围。该扩展带来小尺度、密集、相似实例等挑战,包括关系推理中视觉锚点缺失、跨帧融合时的实例混淆以及动态视角下的归属不确定性。为此,我们提出一种开放词汇管道:利用2D视觉证据锚定细粒度功能边,并通过多线索在3D中跨帧关联节点。同时,将边关联建模为时间图优化,整合证据累积、熵正则化与时间平滑,以稳健确定各节点的功能连接。最后进行全局层次结构重构。大量实验表明,该方法可在真实复杂场景中可靠推断功能性3D场景图,进一步释放其在实际应用中的潜力。

原文摘要 · Abstract (English)

Functional 3D scene graphs offer a versatile and flexible representation for 3D scene understanding and robotic manipulation, defined by object nodes, interactive elements, and functional relationship edges. However, their potential remains underexplored due to the limited coverage of existing benchmarks and the overly straightforward design of previous pipelines, which primarily focus on large-scale furniture but lack of hierarchical structures. Therefore, in this work, we extend the benchmark coverage by introducing dense tabletop objects and explicit multi-level functional relationships. This expansion introduces critical challenges involving small-scale, dense, and similar instances, with lack of visual anchoring in relational reasoning, instance confusion during cross-frame fusion, and attribution uncertainty under dynamic viewpoints. To address these issues, we propose an open-vocabulary pipeline based on 2D visual grounding and 3D graph optimization. Specifically, we anchor fine-grained functional edges from 2D visual evidence, and associate nodes across frames in 3D using multiple cues. Furthermore, edge association is formulated as temporal graph optimization, integrating evidence accumulation, entropy regularization, and temporal smoothing to robustly determine the functional connections of each node. Finally, global hierarchy shaping is performed to recover the hierarchical graph structure. Extensive experiments demonstrate that the proposed method can reliably infer functional 3D scene graphs in challenging real-world scenes, thereby further unlocking their potential for practical applications.

3D场景图功能理解机器人交互视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。