让机器人理解复杂空间指令,如‘找桌上的遥控器’
DIV-Nav: Open-Vocabulary Spatial Relationships for Multi-Object Navigation
- 将复杂空间指令分解为简单物体查询,再融合语义地图
- 通过交集定位多物共存区域,准确率超90%
- 适合需要精准空间定位的机器人导航场景
开放词汇语义地图与物体导航的进步使机器人能够对任意物体进行智能搜索。然而,现有方法通常仅支持如‘电视’或‘蓝地毯’这类简单命名查询。本文提出DIV-Nav,一种实时导航系统,可处理包含空间关系的自由文本查询,例如‘找桌上的遥控器’,同时保持语义地图的鲁棒性。该系统通过三步实现:一、将复杂自然语言指令分解为语义地图上的对象级查询;二、计算各对象语义信念图的交集,定位所有物体共存区域;三、利用大型视觉语言模型(LVLM)验证候选区域是否满足原始空间约束。此外,我们研究了如何调整在线语义地图的前沿探索目标,以更高效引导空间搜索过程。在MultiON基准和波士顿动力Spot机器人上基于Jetson Orin AGX的实机部署中进行了充分验证。
原文摘要 · Abstract (English)
Advances in open-vocabulary semantic mapping and object navigation have enabled robots to perform an informed search of their environment for an arbitrary object. However, such zero-shot object navigation is typically designed for simple queries with an object name like "television" or "blue rug". Here, we consider more complex free-text queries with spatial relationships, such as "find the remote on the table" while still leveraging robustness of a semantic map. We present DIV-Nav, a real-time navigation system that efficiently addresses this problem through a series of relaxations: i) Decomposing natural language instructions with complex spatial constraints into simpler object-level queries on a semantic map, ii) computing the Intersection of individual semantic belief maps to identify regions where all objects co-exist, and iii) Validating the discovered objects against the original, complex spatial constrains via a LVLM. We further investigate how to adapt the frontier exploration objectives of online semantic mapping to such spatial search queries to more effectively guide the search process. We validate our system through extensive experiments on the MultiON benchmark and real-world deployment on a Boston Dynamics Spot robot using a Jetson Orin AGX. More details and videos are available at https://anonsub42.github.io/reponame/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。