实时构建带语义的3D物体级地图,支持零样本理解新物体。
OpenVox: Real-time Instance-level Open-vocabulary Probabilistic Voxel Representation
- 用语言描述增强实例分割,实现语义推理。
- 跨帧融合中分离关联与更新,抗传感器噪声。
- 适合需要动态理解环境的机器人应用。
近年来,视觉语言模型(VLMs)推动了开放词汇映射的发展,使移动机器人能够同时完成环境重建与高层语义理解。尽管集成物体认知有助于缓解点云特征图中的语义歧义,但在实例级别高效获取丰富语义并实现鲁棒的增量重建仍具挑战。为此,我们提出 OpenVox,一种实时增量式开放词汇概率实例体素表示方法。前端设计了高效的实例分割与理解流水线,通过编码描述文本增强语言推理能力;后端采用概率实例体素,并将跨帧增量融合分解为实例关联与地图实时演化两个子任务,提升对传感器和分割噪声的鲁棒性。在多个数据集上的广泛评估表明,OpenVox 在零样本实例分割、语义分割和开放词汇检索任务上达到当前最优性能。真实机器人实验进一步验证了其稳定、实时运行的能力。
原文摘要 · Abstract (English)
In recent years, vision-language models (VLMs) have advanced open-vocabulary mapping, enabling mobile robots to simultaneously achieve environmental reconstruction and high-level semantic understanding. While integrated object cognition helps mitigate semantic ambiguity in point-wise feature maps, efficiently obtaining rich semantic understanding and robust incremental reconstruction at the instance-level remains challenging. To address these challenges, we introduce OpenVox, a real-time incremental open-vocabulary probabilistic instance voxel representation. In the front-end, we design an efficient instance segmentation and comprehension pipeline that enhances language reasoning through encoding captions. In the back-end, we implement probabilistic instance voxels and formulate the cross-frame incremental fusion process into two subtasks: instance association and live map evolution, ensuring robustness to sensor and segmentation noise. Extensive evaluations across multiple datasets demonstrate that OpenVox achieves state-of-the-art performance in zero-shot instance segmentation, semantic segmentation, and open-vocabulary retrieval. Furthermore, real-world robotics experiments validate OpenVox's capability for stable, real-time operation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。