arXiv:2601.08665cs.ROcs.CV2026-01被引 13

让机器人像人一样思考导航,该模型能自动决定何时思考、记住什么。

VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory

  • 引入自适应思维链,按需触发推理,兼顾快速执行与深度规划。
  • 构建视觉辅助语言记忆,实现跨模态长期记忆,避免重复探索。
  • 适合需要长程规划与复杂环境适应的机器人导航任务。

VLA模型在具身导航中展现出巨大潜力,通过统一感知与规划,并继承大模型的强大泛化能力。然而,现有多数VLA模型依赖从观察到动作的直接反应式映射,缺乏显式推理能力和持久记忆,难以应对复杂、长时程的导航任务。为此,我们提出VLingNav,一种基于语言驱动认知的具身导航VLA模型。首先,受人类认知双过程理论启发,引入自适应思维链机制,仅在必要时动态触发显式推理,使智能体可在快速直觉执行与慢速理性规划间自由切换。其次,为处理长时程空间依赖,设计视觉辅助语言记忆模块,构建持久的跨模态语义记忆,使智能体能回溯过往观测,防止重复探索,并推断动态环境中的运动趋势。训练方面,我们构建了当前最大规模的具身导航数据集Nav-AdaCoT-2.9M,包含自适应思维链标注,引导模型学习何时思考、思考什么。此外,引入在线专家指导强化学习阶段,使模型超越纯模仿学习,获得更鲁棒、自我探索的导航行为。大量实验表明,VLingNav在多种具身导航基准上达到领先性能。值得注意的是,该模型可零样本迁移至真实机器人平台,完成多样化导航任务,展现出强大的跨领域与跨任务泛化能力。

原文摘要 · Abstract (English)

VLA models have shown promising potential in embodied navigation by unifying perception and planning while inheriting the strong generalization abilities of large VLMs. However, most existing VLA models rely on reactive mappings directly from observations to actions, lacking the explicit reasoning capabilities and persistent memory required for complex, long-horizon navigation tasks. To address these challenges, we propose VLingNav, a VLA model for embodied navigation grounded in linguistic-driven cognition. First, inspired by the dual-process theory of human cognition, we introduce an adaptive chain-of-thought mechanism, which dynamically triggers explicit reasoning only when necessary, enabling the agent to fluidly switch between fast, intuitive execution and slow, deliberate planning. Second, to handle long-horizon spatial dependencies, we develop a visual-assisted linguistic memory module that constructs a persistent, cross-modal semantic memory, enabling the agent to recall past observations to prevent repetitive exploration and infer movement trends for dynamic environments. For the training recipe, we construct Nav-AdaCoT-2.9M, the largest embodied navigation dataset with reasoning annotations to date, enriched with adaptive CoT annotations that induce a reasoning paradigm capable of adjusting both when to think and what to think about. Moreover, we incorporate an online expert-guided reinforcement learning stage, enabling the model to surpass pure imitation learning and to acquire more robust, self-explored navigation behaviors. Extensive experiments demonstrate that VLingNav achieves state-of-the-art performance across a wide range of embodied navigation benchmarks. Notably, VLingNav transfers to real-world robotic platforms in a zero-shot manner, executing various navigation tasks and demonstrating strong cross-domain and cross-task generalization.

具身导航思维链多模态记忆机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。