一个模型搞定五种机器人导航任务,实现统一规划与执行。
ABot-N0: Technical Report on the VLA Foundation Model for Versatile Embodied Navigation
- 用大模型推理+连续动作生成的分层架构,统一处理多种导航任务。
- 在7个基准上达新最优,使用1690万条轨迹和500万条推理数据训练。
- 适合做复杂环境下的长时序自主导航系统研究者参考。
具身导航长期受制于任务专用模型的碎片化问题。本文提出ABot-N0,一种统一的视觉-语言-动作(VLA)基础模型,实现了五大核心任务的“大一统”:点目标导航、物体目标导航、指令跟随、兴趣点导航和人跟踪。ABot-N0采用分层“脑-行动”架构,由基于大语言模型的认知脑负责语义推理,结合基于流匹配的动作专家实现精确连续轨迹生成。为支持大规模学习,我们构建了ABot-N0数据引擎,整合了1690万条专家轨迹和500万条推理样本,覆盖7802个高保真3D场景(总计10.7 km²)。ABot-N0在7个基准测试中均达到新SOTA性能,显著优于各类专用模型。此外,我们的智能体导航系统融合规划器与分层拓扑记忆,可在动态真实环境中实现鲁棒的长时序任务执行。
原文摘要 · Abstract (English)
Embodied navigation has long been fragmented by task-specific architectures. We introduce ABot-N0, a unified Vision-Language-Action (VLA) foundation model that achieves a ``Grand Unification'' across 5 core tasks: Point-Goal, Object-Goal, Instruction-Following, POI-Goal, and Person-Following. ABot-N0 utilizes a hierarchical ``Brain-Action'' architecture, pairing an LLM-based Cognitive Brain for semantic reasoning with a Flow Matching-based Action Expert for precise, continuous trajectory generation. To support large-scale learning, we developed the ABot-N0 Data Engine, curating 16.9M expert trajectories and 5.0M reasoning samples across 7,802 high-fidelity 3D scenes (10.7 $\text{km}^2$). ABot-N0 achieves new SOTA performance across 7 benchmarks, significantly outperforming specialized models. Furthermore, our Agentic Navigation System integrates a planner with hierarchical topological memory, enabling robust, long-horizon missions in dynamic real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。