arXiv:2510.21069cs.CVcs.RO2025-10被引 1

零样本构建可增量更新的3D场景图,支持机器人环境理解

ZING-3D: Zero-shot Incremental 3D Scene Graphs via Vision-Language Models

  • 用视觉语言模型生成2D语义图,结合深度信息实现3D空间对齐
  • 在Replica和HM3D数据集上无需微调即可准确捕捉物体空间关系
  • 适合需要动态感知与语义理解的机器人应用

理解复杂3D环境需要结构化的场景表示,不仅包含物体,还需其语义与空间关系。现有3D场景图生成方法多限于单视角,无法支持新观测的增量更新,且缺乏显式几何定位,难以适用于具身智能场景。本文提出ZING-3D框架,利用预训练基础模型实现开放词汇识别,在零样本条件下生成丰富的语义场景表示,并支持增量更新与3D空间几何对齐,适用于下游机器人任务。该方法通过视觉语言模型推理生成丰富2D场景图,再利用深度信息将其锚定至3D空间。节点包含开放词汇物体的特征、3D位置与语义上下文,边则表示物体间的空间与语义关系及距离。在Replica与HM3D数据集上的实验表明,ZING-3D无需任务特定训练即可有效捕捉空间与关系知识。

原文摘要 · Abstract (English)

Understanding and reasoning about complex 3D environments requires structured scene representations that capture not only objects but also their semantic and spatial relationships. While recent works on 3D scene graph generation have leveraged pretrained VLMs without task-specific fine-tuning, they are largely confined to single-view settings, fail to support incremental updates as new observations arrive and lack explicit geometric grounding in 3D space, all of which are essential for embodied scenarios. In this paper, we propose, ZING-3D, a framework that leverages the vast knowledge of pretrained foundation models to enable open-vocabulary recognition and generate a rich semantic representation of the scene in a zero-shot manner while also enabling incremental updates and geometric grounding in 3D space, making it suitable for downstream robotics applications. Our approach leverages VLM reasoning to generate a rich 2D scene graph, which is grounded in 3D using depth information. Nodes represent open-vocabulary objects with features, 3D locations, and semantic context, while edges capture spatial and semantic relations with inter-object distances. Our experiments on scenes from the Replica and HM3D dataset show that ZING-3D is effective at capturing spatial and relational knowledge without the need of task-specific training.

3D场景图视觉语言模型零样本机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。