arXiv:2503.19199cs.CV2025-03CVPR被引 62

构建可理解家居功能的3D场景图,让机器懂物品怎么用。

Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces

  • 用视觉语言模型和大语言模型挖掘物品功能知识
  • 在两个新标注数据集上显著超越现有方法
  • 适合做智能机器人、室内问答系统的研究人员

我们提出从带位姿的RGB-D图像中预测真实室内环境的功能性3D场景图。与传统关注物体空间关系的3D场景图不同,功能性3D场景图捕捉物体、交互元素及其功能关系。由于缺乏训练数据,我们利用视觉语言模型(VLMs)和大语言模型(LLMs)编码功能知识。我们在扩展的SceneFun3D数据集和新收集的FunGraph3D数据集上评估方法,两者均标注有功能性3D场景图。所提方法显著优于适配的基线模型(包括Open3DSG和ConceptGraph),证明其在建模复杂场景功能上的有效性。我们还展示了下游应用,如3D问答和机器人操作。

原文摘要 · Abstract (English)

We introduce the task of predicting functional 3D scene graphs for real-world indoor environments from posed RGB-D images. Unlike traditional 3D scene graphs that focus on spatial relationships of objects, functional 3D scene graphs capture objects, interactive elements, and their functional relationships. Due to the lack of training data, we leverage foundation models, including visual language models (VLMs) and large language models (LLMs), to encode functional knowledge. We evaluate our approach on an extended SceneFun3D dataset and a newly collected dataset, FunGraph3D, both annotated with functional 3D scene graphs. Our method significantly outperforms adapted baselines, including Open3DSG and ConceptGraph, demonstrating its effectiveness in modeling complex scene functionalities. We also demonstrate downstream applications such as 3D question answering and robotic manipulation using functional 3D scene graphs. See our project page at https://openfungraph.github.io

3D场景图功能理解机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。