arXiv:2506.09839cs.CVcs.AI2025-06被引 39

构建通用具身导航系统,支持多模态自由指令理解与思考决策。

OctoNav: Towards Generalist Embodied Navigation

  • 基于大模型设计跨任务通用导航框架,支持多模态复合指令。
  • 在连续环境中构建大规模标注数据集,包含动作前思考过程。
  • 采用三阶段训练策略,提升模型推理与导航能力,适合通用智能体研究者。

具身导航是实现通用具身智能的基础。以往研究分属于不同任务(如ObjNav、ImgNav、VLN),任务目标与模态各异,导致数据集和方法孤立。本文提出通用导航基准OctoNav-Bench与模型OctoNav-R1,构建连续环境下的大规模指令-轨迹对,指令涵盖任意多模态组合与能力需求,并引入思考前行动(TBA-CoT)数据集以捕捉行为背后的推理过程。OctoNav-R1基于多模态大模型(MLLM)构建,适配为视觉语言动作(VLA)模型,仅依赖2D视觉观测生成低层动作。设计混合训练范式(HTP),包含三阶段:动作/思考监督微调(TBA-SFT)、导航-广义策略优化(Nav-GPRO)与在线强化学习(Online RL)。其中TBA-SFT与Nav-GPRO受OpenAI-o1与DeepSeek-R1启发,通过思考前行动机制提升模型推理能力。实验表明,OctoNav-R1在多种导航任务上优于现有方法。

原文摘要 · Abstract (English)

Embodied navigation stands as a foundation pillar within the broader pursuit of embodied AI. However, previous navigation research is divided into different tasks/capabilities, e.g., ObjNav, ImgNav and VLN, where they differ in task objectives and modalities, making datasets and methods are designed individually. In this work, we take steps toward generalist navigation agents, which can follow free-form instructions that include arbitrary compounds of multi-modal and multi-capability. To achieve this, we propose a large-scale benchmark and corresponding method, termed OctoNav-Bench and OctoNav-R1. Specifically, OctoNav-Bench features continuous environments and is constructed via a designed annotation pipeline. We thoroughly craft instruction-trajectory pairs, where instructions are diverse in free-form with arbitrary modality and capability. Also, we construct a Think-Before-Action (TBA-CoT) dataset within OctoNav-Bench to provide the thinking process behind actions. For OctoNav-R1, we build it upon MLLMs and adapt it to a VLA-type model, which can produce low-level actions solely based on 2D visual observations. Moreover, we design a Hybrid Training Paradigm (HTP) that consists of three stages, i.e., Action-/TBA-SFT, Nav-GPRO, and Online RL stages. Each stage contains specifically designed learning policies and rewards. Importantly, for TBA-SFT and Nav-GRPO designs, we are inspired by the OpenAI-o1 and DeepSeek-R1, which show impressive reasoning ability via thinking-before-answer. Thus, we aim to investigate how to achieve thinking-before-action in the embodied navigation field, to improve model's reasoning ability toward generalists. Specifically, we propose TBA-SFT to utilize the TBA-CoT dataset to fine-tune the model as a cold-start phrase and then leverage Nav-GPRO to improve its thinking ability. Finally, OctoNav-R1 shows superior performance compared with previous methods.

具身智能导航多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。