arXiv:2505.15373cs.CVcs.RO2025-05被引 5

零样本实现3D开放词汇全景重建,支持自然语言查询

RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation

  • 通过时空聚合融合2D视觉语言模型与3D几何重建
  • 无需训练,实时处理并保持语义一致性
  • 适合需要动态理解新物体的机器人导航场景

构建和理解复杂三维环境是自主系统感知与交互物理世界的基础,需兼具精确的几何重建与丰富的语义理解。现有3D语义映射系统虽能精准识别预定义物体实例,却难以在在线运行时高效构建开放词汇语义地图。尽管近期视觉-语言模型已在2D图像中实现开放词汇物体识别,但尚未跨越至3D空间理解。核心挑战在于开发一个无训练的统一系统,能够实时同步构建高精度3D地图、保持语义一致性,并支持自然语言交互。本文提出RAZER框架,通过在线实例级语义嵌入融合,将GPU加速的几何重建与开放词汇视觉-语言模型无缝结合,基于分层对象关联与空间索引进行引导。该无训练系统通过增量处理和统一的几何-语义更新,实现卓越性能,且对2D分割不一致具有鲁棒性。所提通用3D场景理解框架可应用于零样本3D实例检索、分割与物体检测,以推理未见物体并解析自然语言查询。

原文摘要 · Abstract (English)

Mapping and understanding complex 3D environments is fundamental to how autonomous systems perceive and interact with the physical world, requiring both precise geometric reconstruction and rich semantic comprehension. While existing 3D semantic mapping systems excel at reconstructing and identifying predefined object instances, they lack the flexibility to efficiently build semantic maps with open-vocabulary during online operation. Although recent vision-language models have enabled open-vocabulary object recognition in 2D images, they haven't yet bridged the gap to 3D spatial understanding. The critical challenge lies in developing a training-free unified system that can simultaneously construct accurate 3D maps while maintaining semantic consistency and supporting natural language interactions in real time. In this paper, we develop a zero-shot framework that seamlessly integrates GPU-accelerated geometric reconstruction with open-vocabulary vision-language models through online instance-level semantic embedding fusion, guided by hierarchical object association with spatial indexing. Our training-free system achieves superior performance through incremental processing and unified geometric-semantic updates, while robustly handling 2D segmentation inconsistencies. The proposed general-purpose 3D scene understanding framework can be used for various tasks including zero-shot 3D instance retrieval, segmentation, and object detection to reason about previously unseen objects and interpret natural language queries. The project page is available at https://razer-3d.github.io.

3D重建开放词汇零样本语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。