arXiv:2507.22052cs.CV2025-07被引 12

用视觉视频实现开放词汇语义3D重建,让机器理解空间中的物体含义。

Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos

论文配图:Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
图 1 · 摘自论文原文
  • 融合CLIP语义的3D重建模块,直接在重建中注入物体级语义
  • 2D-3D联合建模,实现细粒度语义与几何的一致对齐
  • 支持开放词汇,适合实时语义感知空间智能应用

我们提出Ov3R,一种从RGB视频流进行开放词汇语义3D重建的新框架,旨在推动空间人工智能的发展。系统包含两个核心组件:CLIP3R,一个受CLIP启发的3D重建模块,能从重叠视频片段中预测稠密点云图,并嵌入物体级语义;以及2D-3D OVS,一个2D-3D开放词汇语义模块,通过学习融合空间、几何与语义线索的特征描述子,将2D特征提升至3D。与以往方法不同,Ov3R将CLIP语义直接融入重建过程,实现了全局一致的几何结构与精细的语义对齐。该框架在稠密3D重建和开放词汇3D分割任务上均达到当前最优性能,标志着向实时、语义感知的空间人工智能迈进一步。

原文摘要 · Abstract (English)

We present Ov3R, a novel framework for open-vocabulary semantic 3D reconstruction from RGB video streams, designed to advance Spatial AI. The system features two key components: CLIP3R, a CLIP-informed 3D reconstruction module that predicts dense point maps from overlapping clips while embedding object-level semantics; and 2D-3D OVS, a 2D-3D open-vocabulary semantic module that lifts 2D features into 3D by learning fused descriptors integrating spatial, geometric, and semantic cues. Unlike prior methods, Ov3R incorporates CLIP semantics directly into the reconstruction process, enabling globally consistent geometry and fine-grained semantic alignment. Our framework achieves state-of-the-art performance in both dense 3D reconstruction and open-vocabulary 3D segmentation, marking a step forward toward real-time, semantics-aware Spatial AI.

3D重建语义理解开放词汇视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。