arXiv:2607.22226cs.RO2026-07

让机器人离线导航更准更快,靠小模型+几何定位

Offline Vision-Language Navigation with Geometric Goal Localization for Outdoor Environments

  • 用17个小型语言模型对比测试,选出发音快又准的本地化方案
  • 结合视觉与激光雷达,把目标定位误差从2.05米降到0.20米
  • 全程无需联网,适合户外无人区智能导航,尤其适合边缘设备

基于基础模型的视觉-语言导航(VLN)通过理解自然语言指令、识别语义目标并遵循用户行为规则,推动了自主机器人导航的发展。然而,现有系统严重依赖云端大模型进行语言理解与语义定位,在无网络或需精确度量目标定位的场景下受限。尽管小型语言模型(SLMs)支持全机载推理,但其在导航指令解析中的适用性尚未系统评估。本文针对室外环境的全机载VLN提出三项贡献:首先,构建首个对17个边缘部署型SLMs与4个在线API的系统性基准测试,涵盖三类计算平台,评估指令分解的准确率与延迟,为本地语言模型选型提供实践指导;其次,提出轻量级混合语义-几何目标定位框架,融合开放词汇目标检测、提示分割与LiDAR几何信息,实现高精度度量目标估计,并在几何观测不可靠时保持视觉方位引导;第三,将上述进展集成至Edge-BehAV,一个完全离线的BehAV架构扩展,实现云无关的行为引导导航。实验表明,最优离线SLM在指令分解性能上媲美最强云端API,速度提升约9倍且无需网络连接;所提定位框架将平均目标距离误差由2.05米降至0.20米,计算成本更低;完整系统在32次闭环室外试验中成功31次。

原文摘要 · Abstract (English)

Foundation-model-based vision-language navigation (VLN) has advanced autonomous robot navigation by enabling robots to interpret natural-language instructions, identify semantic goals, and follow user-specified behavioral rules. However, existing VLN systems rely heavily on cloud-hosted foundation models for language understanding and semantic grounding, limiting their applicability where network connectivity is unavailable and reliable metric goal localization is required. Although recent small language models (SLMs) enable fully onboard inference, their suitability for navigation instruction decomposition has not been systematically evaluated. This paper makes three contributions toward fully onboard VLN for outdoor environments. First, we present the first systematic benchmark of 17 edge-deployable SLMs against 4 online APIs for robotic navigation instruction decomposition, evaluating accuracy and latency on human-annotated instructions across three computing platforms and providing practical guidance for selecting onboard language models. Second, we propose a lightweight hybrid semantic-geometric goal localization framework that combines open-vocabulary object detection, prompted segmentation, and LiDAR geometry to estimate metric goals, while maintaining visual bearing guidance when reliable geometric observations are unavailable. Third, we integrate these advances into Edge-BehAV, a fully onboard extension of the BehAV architecture that enables cloud-independent behavior-guided navigation. Experimental results show that the best offline SLM matches the instruction decomposition performance of the strongest cloud API while running approximately 9x faster and without network connectivity. The proposed goal localization framework reduces mean goal-distance error from 2.05 m to 0.20 m at lower computational cost, and the complete system succeeds in 31 of 32 closed-loop outdoor trials.

视觉导航小模型离线推理目标定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。