arXiv:2509.13733cs.RO2025-09被引 20

用分层图谱+快慢推理,让机器人导航又快又准

FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph

  • 构建分层多模态场景图,支持从房间到物体的渐进式定位
  • 在4个数据集上成功率超前,响应时间比纯视觉语言模型快82%
  • 适合需要实时导航与自然语言交互的机器人系统

视觉-语言导航(VLN)是机器人系统中的基础挑战,广泛应用于具身智能体在真实环境中的部署。尽管近期取得进展,现有方法在长距离空间推理方面仍受限,常表现出低成功率和高推理延迟,尤其在长程导航任务中。为此,我们提出FSR-VLN,结合分层多模态场景图(HMSG)与快-慢推理机制(FSR)。HMSG提供多模态地图表示,支持从粗粒度房间级定位到细粒度目标视角与物体识别的渐进检索。基于HMSG,FSR先进行快速匹配以高效筛选候选房间、视角和物体,再通过视觉语言模型驱动的精炼完成最终目标选择。我们在四个由人形机器人采集的综合性室内数据集上评估了FSR-VLN,使用87条涵盖多样化物体类别的指令。结果表明,FSR-VLN在所有数据集上均达到当前最优表现,以检索成功率(RSR)衡量;相比基于视觉语言模型的方法,在导览视频上的响应时间减少82%。此外,我们还将FSR-VLN集成至Unitree-G1人形机器人,结合语音交互、规划与控制模块,实现了自然语言交互与实时导航。

原文摘要 · Abstract (English)

Visual-Language Navigation (VLN) is a fundamental challenge in robotic systems, with broad applications for the deployment of embodied agents in real-world environments. Despite recent advances, existing approaches are limited in long-range spatial reasoning, often exhibiting low success rates and high inference latency, particularly in long-range navigation tasks. To address these limitations, we propose FSR-VLN, a vision-language navigation system that combines a Hierarchical Multi-modal Scene Graph (HMSG) with Fast-to-Slow Navigation Reasoning (FSR). The HMSG provides a multi-modal map representation supporting progressive retrieval, from coarse room-level localization to fine-grained goal view and object identification. Building on HMSG, FSR first performs fast matching to efficiently select candidate rooms, views, and objects, then applies VLM-driven refinement for final goal selection. We evaluated FSR-VLN across four comprehensive indoor datasets collected by humanoid robots, utilizing 87 instructions that encompass a diverse range of object categories. FSR-VLN achieves state-of-the-art (SOTA) performance in all datasets, measured by the retrieval success rate (RSR), while reducing the response time by 82% compared to VLM-based methods on tour videos by activating slow reasoning only when fast intuition fails. Furthermore, we integrate FSR-VLN with speech interaction, planning, and control modules on a Unitree-G1 humanoid robot, enabling natural language interaction and real-time navigation.

视觉导航多模态人形机器人推理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。