arXiv:2607.10383cs.CVcs.AI2026-07被引 2

通过分层设计实现视觉语言导航的通用性与可解释性

ABot-N1: Toward a General Visual Language Navigation Foundation Model

论文配图:ABot-N1: Toward a General Visual Language Navigation Foundation Model
图 1 · 摘自论文原文
  • 慢速推理模块生成像素级目标点,快速执行模块实时规划路径
  • 在城市级导航中提升35%到达率,复杂场景成功率超92%
  • 支持多种任务且结果可追溯,适合真实世界机器人应用

视觉语言导航基础模型旨在统一空间决策的深度推理与多样化具身任务的广泛适应性。现有方法多采用端到端策略直接映射观测到动作,但普遍存在坐标漂移和长尾语义处理不佳问题,且黑箱机制缺乏可解释性。本文提出ABot-N1,通过慢-快双架构解耦认知与控制,利用双视觉语言信号引导。慢速视觉语言推理器进行显式思维链推理并生成像素级目标点,作为点目标、物体目标、兴趣点目标、指令跟随和人追踪等任务的通用接口。快速动作专家结合文本线索与像素指引,在原生控制频率下生成连续路径点。该设计通过像素锚点与显式语言轨迹连接高层意图与底层控制,显著提升导航的鲁棒性、泛化性和可解释性。在仿真与真实世界基准上均取得新最优性能:城市级导航中兴趣点到达率提升35.0%(达77.3%),复杂室内外场景成功率分别达95.4%和92.9%;同时在物体抓取、人追踪和指令跟随任务中保持优异鲁棒性。论文开源了新的点目标/兴趣点目标基准,推动城市级导航研究发展。

原文摘要 · Abstract (English)

Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet they often suffer from coordinate drift and poor handling of long-tail semantics. Furthermore, these black-box mappings lack interpretability, hindering the simultaneous achievement of generality, robustness, and transparency. We present ABot-N1, a step toward a general Visual Language Navigation foundation model, that addresses these challenges by decoupling cognition from control via a slow-fast architecture guided by dual visual-language signals. More specifically, a slow vision-language reasoner performs explicit Chain-of-Thought reasoning while producing a pixel goal. This compact set of image-space anchor points serves as a universal interface for diverse tasks, including point-goal, object-goal, poi-goal, instruction-following, and person-following. Subsequently, a fast action expert leverages both the textual cues and the pixel guidance to generate continuous waypoints at the native control frequency. By bridging high-level intents and low-level control through pixel-grounded anchors paired with explicit linguistic traces, our approach ensures robust, generalizable, and interpretable navigation across simulation and real-world benchmarks. ABot-N1 establishes new state-of-the-art records, delivering massive gains specifically in urban-scale navigation: boosting POI arrival by 35.0% (to 77.3%) and achieving 95.4%/92.9% SR in complex indoor and outdoor scenes. It also maintains superior robustness across object-reaching, person-following, and instruction-following tasks. New Point-Goal/POI-Goal benchmarks are released as open source to advance the field of urban-scale navigation.

视觉导航多模态可解释性机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。