arXiv:2506.14233cs.RO2025-06被引 3

让机器人实时理解语言和社交线索,实现更聪明的导航。

Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments

  • 用自监督学习将语言推理嵌入视觉编码器的潜在空间
  • 在未见数据集上提升52.94%,真实场景提升41.67%
  • 适合需要理解人类意图的智能机器人导航任务

大型视觉语言模型(VLMs)在理解上下文线索、人类意图和社会动态方面展现出潜力,可提升人机交互环境中的移动机器人导航能力。然而,其计算复杂度高且对连续数值数据敏感性不足,限制了实时性能与精确运动控制。为此,我们提出Narrate2Nav,一种新型实时视觉-动作模型,采用基于Barlow Twins冗余消除损失的自监督学习框架,将隐式自然语言推理、社会线索和人类意图嵌入视觉编码器中,实现模型潜在空间内的推理而非标记空间。训练时融合RGB输入、运动指令和场景文本信号,使模型从机器人观测直接生成短时程点目标导航的低层运动指令。在离线未见数据集和真实世界实验中,相较于最佳基线分别提升52.94%和41.67%。定性对比显示,其视觉编码器注意力图更聚焦于导航关键场景元素,证明其在人机交互导航任务中的有效性。

原文摘要 · Abstract (English)

Large Vision-Language Models (VLMs) have demonstrated potential in enhancing mobile robot navigation in human-centric environments by understanding contextual cues, human intentions, and social dynamics while exhibiting reasoning capabilities. However, their computational complexity and limited sensitivity to continuous numerical data impede real-time performance and precise motion control. To this end, we propose Narrate2Nav, a novel real-time vision-action model that leverages a novel self-supervised learning framework based on the Barlow Twins redundancy reduction loss to embed implicit natural language reasoning, social cues, and human intentions within a visual encoder-enabling reasoning in the model's latent space rather than token space. The model combines RGB inputs, motion commands, and textual signals of scene context during training to bridge from robot observations to low-level motion commands for short-horizon point-goal navigation during deployment. Extensive evaluation of Narrate2Nav across various challenging scenarios in both offline unseen dataset and real-world experiments demonstrates an overall improvement of 52.94 percent and 41.67 percent, respectively, over the next best baseline. Additionally, qualitative comparative analysis of Narrate2Nav's visual encoder attention map against four other baselines demonstrates enhanced attention to navigation-critical scene elements, underscoring its effectiveness in human-centric navigation tasks.

视觉导航语言推理机器人自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。