实时将视觉语言嵌入映射到精确3D空间,支持自然语言定位物体。
Real-Time 3D Vision-Language Embedding Mapping
- 用局部掩码和置信加权融合提升3D嵌入分布清晰度与可靠性。
- 在多房间与单个物体层面均实现精准语义定位,满足实时性要求。
- 无需额外数据,仅需图像即可用于手持、移动机器人等交互任务。
精确的语义3D表征对众多机器人任务至关重要。本文提出一种简单而高效的方法,实现实时将视觉语言模型的2D嵌入整合到度量准确的3D表示中。通过局部嵌入掩码策略增强嵌入分布区分度,并采用置信加权3D融合提升3D嵌入可靠性。所得的度量准确嵌入表示具有任务无关性,可在全局多房间场景及局部物体层级表达语义概念,从而支持多种需通过自然语言定位感兴趣物体的交互式机器人应用。我们在多个真实世界序列上评估该方法,验证其在提升物体定位准确性的同时,显著改善运行效率以满足实时约束。进一步实验展示了该方法在手持、移动机器人及操作任务中的广泛适用性,仅依赖原始图像数据即可完成。
原文摘要 · Abstract (English)
A metric-accurate semantic 3D representation is essential for many robotic tasks. This work proposes a simple, yet powerful, way to integrate the 2D embeddings of a Vision-Language Model in a metric-accurate 3D representation at real-time. We combine a local embedding masking strategy, for a more distinct embedding distribution, with a confidence-weighted 3D integration for more reliable 3D embeddings. The resulting metric-accurate embedding representation is task-agnostic and can represent semantic concepts on a global multi-room, as well as on a local object-level. This enables a variety of interactive robotic applications that require the localisation of objects-of-interest via natural language. We evaluate our approach on a variety of real-world sequences and demonstrate that these strategies achieve a more accurate object-of-interest localisation while improving the runtime performance in order to meet our real-time constraints. We further demonstrate the versatility of our approach in a variety of interactive handheld, mobile robotics and manipulation tasks, requiring only raw image data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。