实现动态环境下的实时开放词汇实例语义建图,支持零样本识别新物体。
OVI-MAP:Open-Vocabulary Instance-Semantic Mapping
- 分离实例重建与语义推理,仅在关键视图用视觉语言模型提取语义特征。
- 在标准基准上优于现有方法,实现稳定跟踪与零样本语义标注。
- 适合需要长期探索的自动驾驶或机器人导航场景。
增量式开放词汇3D实例语义建图对复杂日常环境中自主代理至关重要。然而,由于需兼顾鲁棒实例分割、实时处理及灵活的开放集推理,仍具挑战性。现有方法常依赖封闭集假设或密集像素级语言融合,限制可扩展性与时间一致性。本文提出OVI-MAP,将实例重建与语义推断解耦:从RGB-D输入增量构建无类别依赖的3D实例地图,同时仅在自动选取的小部分视图中使用视觉语言模型提取语义特征。该设计实现了稳定的实例追踪与全程零样本语义标注。系统可实时运行,在标准基准上超越当前最优开放词汇建图基线。
原文摘要 · Abstract (English)
Incremental open-vocabulary 3D instance-semantic mapping is essential for autonomous agents operating in complex everyday environments. However, it remains challenging due to the need for robust instance segmentation, real-time processing, and flexible open-set reasoning. Existing methods often rely on the closed-set assumption or dense per-pixel language fusion, which limits scalability and temporal consistency. We introduce OVI-MAP that decouples instance reconstruction from semantic inference. We propose to build a class-agnostic 3D instance map that is incrementally constructed from RGB-D input, while semantic features are extracted only from a small set of automatically selected views using vision-language models. This design enables stable instance tracking and zero-shot semantic labeling throughout online exploration. Our system operates in real time and outperforms state-of-the-art open-vocabulary mapping baselines on standard benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。