arXiv:2509.12129cs.RO2025-09被引 74

跨机器人类型与任务的导航基础模型,无需微调即可通用

Embodied Navigation Foundation Model

  • 统一架构处理多视角、多时序的导航输入
  • 在八百万样本上训练,支持四足、无人机等八类实体
  • 适合需要强泛化能力的现实场景部署

导航是具身智能的核心能力,需根据语言指令感知并交互物理环境。尽管大视觉-语言模型在通用视觉-语言任务中表现优异,其在具身导航中的泛化能力仍受限于狭窄任务和特定构型。本文提出跨实体、跨任务的导航基础模型(NavFoM),基于八百万条包含四足机器人、无人机、轮式机器人及车辆的导航数据,覆盖视觉语言导航、目标搜索、目标追踪和自动驾驶等多样化任务。NavFoM采用统一架构,处理不同相机配置与导航视野的多模态输入,并引入标识符令牌编码相机视角与任务时间上下文。为适应实际部署限制,在有限令牌长度下采用动态采样策略控制观测信息。在多个公开基准测试中,模型无需任务微调即达顶尖或竞争力表现。真实世界实验进一步验证了其强大泛化性与实用性。

原文摘要 · Abstract (English)

Navigation is a fundamental capability in embodied AI, representing the intelligence required to perceive and interact within physical environments following language instructions. Despite significant progress in large Vision-Language Models (VLMs), which exhibit remarkable zero-shot performance on general vision-language tasks, their generalization ability in embodied navigation remains largely confined to narrow task settings and embodiment-specific architectures. In this work, we introduce a cross-embodiment and cross-task Navigation Foundation Model (NavFoM), trained on eight million navigation samples that encompass quadrupeds, drones, wheeled robots, and vehicles, and spanning diverse tasks such as vision-and-language navigation, object searching, target tracking, and autonomous driving. NavFoM employs a unified architecture that processes multimodal navigation inputs from varying camera configurations and navigation horizons. To accommodate diverse camera setups and temporal horizons, NavFoM incorporates identifier tokens that embed camera view information of embodiments and the temporal context of tasks. Furthermore, to meet the demands of real-world deployment, NavFoM controls all observation tokens using a dynamically adjusted sampling strategy under a limited token length budget. Extensive evaluations on public benchmarks demonstrate that our model achieves state-of-the-art or highly competitive performance across multiple navigation tasks and embodiments without requiring task-specific fine-tuning. Additional real-world experiments further confirm the strong generalization capability and practical applicability of our approach.

具身导航基础模型多模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。