用预训练视觉语言模型实现通用机器人导航,无需特定任务头。
LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

- 通过统一的指代标记和动作量化器,让大模型直接理解空间意图并生成动作。
- 在10个公开仿真环境中达到最先进单目导航成功率,覆盖2000+场景、4000+小时数据。
- 零样本跨机器人、场景和目标类型泛化,适合通用导航系统研究者。
具身导航要求智能体将异构目标与视觉观测转化为跨任务、环境及机器人形态的动作。现代视觉语言模型(VLM)已具备视觉定位、空间推理与指向的空间先验,但这些能力极少被直接用于机器人控制。现有导航系统依赖任务或形态特定组件,割裂感知、推理与行动,泛化能力有限。本文提出LightNav-0,一个轻量级通用具身导航模型,通过激发预训练VLM的空间智能并将其对齐导航任务,无需任务特定预测头。LightNav-0以统一标记接口表示多样导航任务:双通道指向表达任务、场景与形态无关的空间意图,残差向量量化的动作分词器将该意图映射为精确的形态相关轨迹。结合时间感知视觉历史压缩、ER中段训练、监督微调与强化学习,该框架支持指令跟随、开放词汇物体导航与视觉跟踪。导航训练语料涵盖2000+场景与4000+小时具身导航数据。用于初始化LightNav-0的LightNav-ER模型,在8个具身推理基准上取得最高完整集平均表现;LightNav-0在全部10个公开导航仿真设置中达到当前最优单目成功率达。真实世界评估进一步证明其在不同机器人形态、多样化场景及静态与动态目标上的零样本泛化能力。结果确立了轻量级VLM作为通用具身导航统一且可迁移的骨干架构。
原文摘要 · Abstract (English)
Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial priors for visual grounding, spatial reasoning, and pointing, but these capabilities are rarely elicited directly for robot control. Existing navigation systems instead rely on task- or embodiment-specific components, fragmenting perception, reasoning, and action while offering limited generalization. Here we present LightNav-0, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads. LightNav-0 represents diverse navigation tasks through a unified token interface: dual-channel pointing expresses task-, scene-, and embodiment-agnostic spatial intent, while a residual vector-quantized action tokenizer maps this intent to precise, embodiment-specific trajectories. Together with temporally aware visual history compression, ER mid-training, supervised fine-tuning, and reinforcement learning, this formulation supports instruction following, open-vocabulary object navigation, and visual tracking within a single model. The navigation training corpus spans 2K+ scenes and 4K+ hours of embodied navigation data. LightNav-ER, the embodied-reasoning checkpoint used to initialize LightNav-0, attains the highest complete-set average across 8 embodied-reasoning benchmarks, while LightNav-0 achieves state-of-the-art monocular success rates across all 10 public navigation simulation settings. Real-world evaluations further demonstrate zero-shot generalization across robot embodiments, diverse scenes, and static and dynamic targets. These results establish compact VLMs as a unified and transferable backbone for generalist embodied navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。