arXiv:2502.00931cs.ROcs.CV2025-02被引 11

用神经符号推理让机器人听懂复杂指令,高效导航不迷路。

VL-Nav: Neuro-Symbolic Reasoning-based Vision-Language Navigation

  • 结合符号系统与神经网络,分解任务并动态规划路径。
  • 室内成功率83.4%,室外75%,真实场景483米长跑成功。
  • 适合复杂指令、多层环境下的智能导航研究者使用。

基于复杂抽象人类指令在未见的大规模环境中导航,仍是自主移动机器人的重大挑战。这要求机器人能推断隐含语义,并高效探索大规模任务空间。然而,现有方法从端到端学习到基于基础模型的模块化架构,往往缺乏任务分解能力或高效的探索策略,导致机器人盲目游荡或目标识别失败。为此,我们提出VL-Nav,一种神经符号(NeSy)视觉语言导航系统。该系统通过两个核心组件实现神经推理与符号引导的融合:(1) 神经符号任务规划器,利用符号化3D场景图与图像记忆系统,增强视觉语言模型(VLMs)的神经推理能力,以完成任务分解与重规划;(2) 神经符号探索系统,将神经语义线索与符号启发函数结合,高效获取任务相关资讯,同时最小化重复行程。在DARPA TIAMAT挑战任务上验证,系统在室内环境达到83.4%成功率,在室外场景达75%。真实实验中,系统取得86.3%成功率,包含一次483米的挑战性长距运行。最后,我们在3D多楼层场景中验证了其对复杂指令的处理能力。

原文摘要 · Abstract (English)

Navigating unseen, large-scale environments based on complex and abstract human instructions remains a formidable challenge for autonomous mobile robots. Addressing this requires robots to infer implicit semantics and efficiently explore large-scale task spaces. However, existing methods, ranging from end-to-end learning to foundation model-based modular architectures, often lack the capability to decompose complex tasks or employ efficient exploration strategies, leading to robot aimless wandering or target recognition failures. To address these limitations, we propose VL-Nav, a neuro-symbolic (NeSy) vision-language navigation system. The proposed system intertwines neural reasoning with symbolic guidance through two core components: (1) a NeSy task planner that leverages a symbolic 3D scene graph and image memory system to enhance the vision language models' (VLMs) neural reasoning capabilities for task decomposition and replanning; and (2) a NeSy exploration system that couples neural semantic cues with the symbolic heuristic function to efficiently gather the task-related information while minimizing unnecessary repeat travel during exploration. Validated on the DARPA TIAMAT Challenge navigation tasks, our system achieved an 83.4% success rate (SR) in indoor environments and 75% in outdoor scenarios. VL-Nav achieved an 86.3% SR in real-world experiments, including a challenging 483-meter run. Finally, we validate the system with complex instructions in a 3D multi-floor scenario.

视觉语言导航神经符号推理机器人导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。