让机器人在未知环境中零样本识别3D物体,提升地图的智能感知能力
OV-MAP: Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
- 用2D掩码投影+合成深度图,实现无需3D标注的3D实例分割
- 在ScanNet200和Replica数据集上零样本性能领先,跨环境适应性强
- 适合需要快速部署、无标注数据支持的机器人导航与交互场景
我们提出OV-MAP,一种面向移动机器人的开放世界3D建图新方法,通过将开放特征融入3D地图以增强物体识别能力。当相邻体素的特征重叠时,特征会溢出体素边界,导致实例级精度下降。为此,我们采用类无关分割模型将2D掩码投影至3D空间,并融合原始点云与合成深度图生成补充深度图像。结合3D掩码投票机制,该方法无需依赖3D监督分割模型即可实现精确的零样本3D实例分割。我们在ScanNet200和Replica等公开数据集上进行充分实验,验证了方法在零样本性能、鲁棒性及多环境适应性方面的优势。此外,真实场景实验进一步证明了其在复杂现实环境中的可扩展性与稳定性。
原文摘要 · Abstract (English)
We introduce OV-MAP, a novel approach to open-world 3D mapping for mobile robots by integrating open-features into 3D maps to enhance object recognition capabilities. A significant challenge arises when overlapping features from adjacent voxels reduce instance-level precision, as features spill over voxel boundaries, blending neighboring regions together. Our method overcomes this by employing a class-agnostic segmentation model to project 2D masks into 3D space, combined with a supplemented depth image created by merging raw and synthetic depth from point clouds. This approach, along with a 3D mask voting mechanism, enables accurate zero-shot 3D instance segmentation without relying on 3D supervised segmentation models. We assess the effectiveness of our method through comprehensive experiments on public datasets such as ScanNet200 and Replica, demonstrating superior zero-shot performance, robustness, and adaptability across diverse environments. Additionally, we conducted real-world experiments to demonstrate our method's adaptability and robustness when applied to diverse real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。